Understanding Autocorrelation
Independent observations each bring their own information. Dependent ones do not: if today is much like yesterday, the second measurement adds less than the first. Autocorrelation quantifies how much less, by correlating the series with itself at a lag, and it is simultaneously the thing that makes forecasting possible and the thing that invalidates the standard formulas.
The simplest generator of such memory is the AR(1) process, in which each value is a fraction phi of the previous one plus fresh noise. Its theoretical autocorrelation at lag k is exactly phi^k. Measured on 39,800 points with phi = 0.7, the autocorrelations at lags 1, 2, 3 and 5 come out as 0.7039, 0.4980, 0.3527 and 0.1796, against a theory of 0.7, 0.49, 0.343 and 0.1681 - agreement to within 0.012. The decay is geometric: a shock is passed along, weakening at each step but never cancelled.
The practical consequence is a specific, computable penalty. The variance of the sample mean of an AR(1) series exceeds the independent-data formula by a factor of (1+phi)/(1-phi), which at phi = 0.7 is 5.6667. Since the standard error is the square root of the variance, it is understated by a factor of about 2.4 when computed as though the points were independent. Intervals are too narrow and tests reject too often, with nothing on the surface to indicate it.
Reading the autocorrelation plot is also how model identification starts: a geometric decay suggests an autoregressive term, a sharp cut-off after a few lags suggests a moving-average one, and a peak at a fixed distance suggests seasonality. Every extra lag is a parameter estimated from a finite and effectively smaller sample, so the honest check on an order chosen this way is forecasting performance on data that comes afterwards.
How to Calculate
ρ_k = Corr(X_t, X_{t+k}); AR(1): ρ_k = φ^k; Var(X̄) inflated by (1+φ)/(1−φ)
where
- ρ_k
- the autocorrelation at lag k: the correlation of the series with itself k steps apart
- φ
- the AR(1) coefficient; stationary when its magnitude is below one
- (1+φ)/(1−φ)
- how much larger the variance of an average is than the independent-data formula
- effective n
- the nominal sample size divided by that inflation factor
Example of Autocorrelation
AR(1) with phi = 0.7 on 39,800 points: autocorrelations of 0.7039, 0.4980, 0.3527 and 0.1796 at lags 1, 2, 3 and 5, against a theoretical phi^k of 0.7, 0.49, 0.343 and 0.1681.
At that phi the variance of the sample mean is 5.6667 times the independent-data value, so a naive standard error is understated by roughly a factor of 2.4 and the effective sample size is about one sixth of the nominal one.
The same dependence is what makes a shuffled validation split misleading: on a dependent series, interpolating a held-out point from the training points around it gives RMSE 0.7515 while honestly forecasting forward gives 4.2074.
Frequently Asked Questions
Is autocorrelation a problem or a resource?
Both, and which one depends on the question. It is what makes the next value predictable from the last, and it is what breaks any procedure assuming independent observations. Forecasting exploits it; inference must account for it.
What if the autocorrelation never seems to decay?
A sample autocorrelation that stays high across many lags usually indicates non-stationarity rather than very long memory. Difference the series and look again; if the pattern disappears, the level was wandering.
Can I just add lagged values as features to a regression?
Often yes, and that is what an autoregressive model is. What does not carry over is the standard-error machinery: the residuals must be checked for remaining autocorrelation, because whatever memory the model failed to capture is still in them.
The Bottom Line
Autocorrelation is the memory of a series, and it decides both what can be forecast and how much any estimate from that series is really worth. Its cost is exact and computable - at phi = 0.7 an average carries about one sixth of the information its sample size suggests.