The Regression That Finds a Relationship That Is Not There
Two series generated from separate random numbers come out significantly related 82.8% of the time, a standard error on dependent data is too narrow by a computable factor of 2.4, and the usual validation split reports a forecaster more than five times better than it is. Three failures, one cause, and the checks that catch each of them.
Prerequisites: Statistical Inference
Write ten lines of code. Generate one series of 200 numbers by starting at zero and adding a random step each time. Generate a second series the same way, from separate random draws, so that by construction neither knows anything about the other. Regress one on the other and test the slope at the usual 5% level.
Do that 2,000 times and the slope comes out significant in 82.8% of them.
A valid test on unrelated data should reject 5% of the time. This one rejects sixteen times too often, and every p-value is computed by the formula that works everywhere else.
That result is the entrance to a subject, and the subject is what happens to ordinary statistics when the rows are in an order that matters.
Failure one: the level that never comes back
Each of those series is a random walk: with independent standard-normal steps. The key property is that it has no level to return to. Whatever value it reaches, the next step starts from there, so displacement accumulates instead of averaging away.
That shows up as a variance which grows with time. Measured over 3,000 replications:
| time | variance of |
|---|---|
| 50 | 48.9 |
| 200 | 194.1 |
| 800 | 757.4 |
Roughly itself, which is the theoretical answer for a sum of independent unit-variance steps.
Now take two such series over any finite window. Each wanders somewhere and stays a while, because leaving is easy and returning is not, so the two will very often be drifting in a consistent direction relative to each other. A scatter of one against the other looks like a line. The regression, built on the assumption of independent draws, counts 200 points of shared drift as 200 independent pieces of evidence when they are closer to one.
The condition being violated is stationarity: the requirement that the statistical behaviour of a series does not depend on when you look at it. In its working form it asks for a mean that does not depend on , a variance that does not depend on , and a covariance that depends on the gap only. A random walk fails the last two.
Why that matters for inference more than for description: every estimate borrows strength across observations, which only makes sense if those observations describe the same process. When the level wanders, an average over the window is an average over different worlds, and an interval around it describes a quantity that never existed.
The standard repair is first differencing. The differences of a random walk are exactly the independent steps it was built from, so the assumption is restored and the same 2,000 pairs now reject at 4.2%, which is what a valid 5% test should produce. The cost should be stated rather than buried: the model now relates changes to changes, which is a different question from the one about levels.
The uncomfortable part is what more data does. The spurious rejection rate does not fall as the series lengthens. It rises, because a longer window gives both series more room to wander. That is the signature of bias rather than noise, and it makes this the same shape of problem as confounding: a property of how the data came to exist, not of how much of it there is.
Failure two: the standard error that is too narrow
Suppose you do have a stationary series. Dependence has not gone away, and it now has a price you can compute exactly.
The simplest process with memory is AR(1): with . Multiply through by and take expectations, and the noise term drops out, leaving autocorrelations that decay geometrically:
Measured on 39,800 simulated points with , the autocorrelations at lags 1, 2, 3 and 5 come out as 0.7039, 0.4980, 0.3527 and 0.1796, against a theory of 0.7, 0.49, 0.343 and 0.1681. Agreement to within 0.012.
Now the consequence. The variance of the sample mean of such a series is not . Summing covariances over all pairs gives an inflation factor that, for the AR(1), collapses to a geometric series:
At that is . Two readings, both worth carrying:
- the standard error is a square root, so computing it as though the points were independent understates it by
- the effective sample size is divided by 5.6667, so 6,000 dependent points carry about as much information about the mean as 1,060 independent ones
Nothing in the output announces this. The estimate is not even biased. It is a right answer reported with far more confidence than it earned, which is why this failure is more common than the dramatic one: it never looks wrong.
The direction checks out at the edges. At the factor is 1, the independent case. As it diverges, which is the random walk having no usable mean at all. Negative gives a factor below 1, because alternating errors partly cancel and such a series carries more information per point than an independent one.
One noise sequence drives the figure below, and the dial only changes how long the series remembers it. Watch two things at once. The autocorrelation bars sit on the theoretical curve, which is what makes the plot a diagnostic rather than a decoration. And the effective sample size falls away while the series above it goes on looking like a series - no warning appears on the surface of the output, which is exactly why this failure is the common one.
Interactive: the same shocks, a longer memory
One noise sequence throughout. Only how long it is remembered changes.
The series
Autocorrelation: measured against φᵏ
- Measured ρ₁
- 0.6970
- Variance inflation
- 5.6667
- Effective n, from 6000
- 1059
- Interval too narrow by
- 2.38x
The measured bars sit on the theoretical φᵏ curve, which is what makes the autocorrelation function a diagnostic rather than a decoration: a geometric decay is the fingerprint of an AR(1), and its rate is φ itself. The variance of the sample mean is inflated by 5.6667, so 6,000 dependent points are worth about 1059 independent ones. Push φ up and watch that number fall while the picture above barely changes.
Failure three: the split that measures the wrong task
Cross-validation is the most reliable habit in applied machine learning, and on a time series it can report a model more than five times better than it is.
Take a dependent series of 240 points, hold out 48, and measure the simplest method for each task - fill a held-out point in from the training points around it, or carry the last value forward:
| split | RMSE |
|---|---|
| 48 points held out at random | 0.7515 |
| the last 48, trained on the first 192 | 4.2074 |
Random splitting is correct when rows are exchangeable - when a row's position carries no information. On a series with memory, position is most of the information. Shuffle, and every held-out point lands between training points, two in three of them with one right beside it on each side, so the model is no longer being asked what happens next. It is being asked to fill a gap between two known values, which on a dependent series is easy.
The 0.75 is therefore not a mildly optimistic estimate of forecasting skill. It is an accurate answer to a question nobody asked.
The rule that prevents this is simple to state - no information from after the forecast origin may reach the model - and it catches more than shuffling. Scaling computed over the whole series lets the training data know the future mean and variance. Imputing a gap from later observations smuggles interpolation into the training set. A centred rolling mean includes the point it is meant to predict. And choosing a model order by looking at the test period spends the test set without recording that it was spent, which is the easiest of the four to commit because it happens across sessions rather than inside one script.
The figure below runs both hold-outs on its own series. It is an independent implementation with an independent generator, and it reads 0.7557 and 4.2224 against the 0.7515 and 4.2074 above, which is the useful thing to notice: the factor of 5.6 belongs to the procedure, not to anyone's seed. Widen the hold-out and watch the two diverge further, because forecast error accumulates with the horizon and gap-filling barely grows.
Interactive: one series under two hold-outs
One series with memory. Only the split changes.
- Shuffled hold-out
- 0.7557
- Forecast forward
- 4.2224
- Flattered by
- 5.6x
- Exact, pooled
- 4.9497
Holding out 48 points at random gives an RMSE of 0.7557; holding out the last 48 gives 4.2224. A factor of 5.6, on the same data. Shuffle a series with memory and every test point sits between training points, two in three with one right beside it on each side, so the model is no longer asked what happens next - it is asked to fill a gap, which is easy here. Drag the hold-out wider: the gap-filling error rises a little as the gaps lengthen, while the forecast error grows far faster, because forecast error accumulates with the horizon.
The baseline that settles arguments
One more measurement, because it costs a line and prevents a category of self-deception. Forecasting 40 points from 120 of history on a random walk:
| forecast | RMSE |
|---|---|
| the last observed value | 4.0419 |
| the mean of the training window | 6.5258 |
| extrapolating the fitted drift | 4.4063 |
The simplest wins. The mean assumes the series returns to a level it does not have. The drift extrapolates a slope that is just the accumulated noise of the training window divided by its length - real in the sample, absent from the process.
So put the naive forecast in every comparison. If a method cannot beat "tomorrow looks like today", its apparent skill is coming from somewhere other than the data, and very often from a split like the one above.
What to actually do
Four habits cover most of it.
- Plot the series before anything else. A wandering level or a widening spread is usually visible. Formal stationarity tests have poor power against slowly moving alternatives, so a test that fails to reject is weak evidence rather than a clearance.
- Difference when the level wanders, and say so. Report that the model describes changes. Do not difference a series that was already stationary - that adds noise for nothing.
- Never report a standard error from dependent data without correcting it. For anything close to AR(1), is the factor, and the effective sample size is the honest number to quote.
- Split forward, at several origins. Every training set ends before its test set begins, and repeating at several cut points shows how performance varies with when the forecast was made. Report the spread, not only the average.
Three failures, one cause. None of them announces itself in the output, and none is fixed by collecting more data. What fixes them is knowing the rows are ordered, and letting that order constrain the test you run, the interval you report, and above all the split you evaluate on.
References & further reading
- Rob J Hyndman, George Athanasopoulos, Forecasting: Principles and Practice, OTexts, 2014source ↗
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.