Skip to content
Kudos AI

Stationarity and the Spurious Regression

Two series with no connection whatever, regressed on each other and found significant 82.8% of the time, the property of the data that causes it, and the one-line transformation that restores an honest test.

IntermediateModule 125 min · 100 XP
Two independently generated random walks drifting apart and back together, a scatter of one against the other tightening into a convincing line, and the same scatter in first differences collapsing into a shapeless cloud.

Here is a result you can reproduce in ten lines, and it should be alarming.

Take two series that have nothing whatever to do with each other - generated from separate random numbers, on separate lines of a script, with no shared input of any kind. Regress one on the other and test the slope at the usual 5% level.

Across 2,000 such pairs of 200 points, the slope comes out significant 82.8% of the time.

Not 5%, which is what a valid test on unrelated data should give. Eighty-three percent. The arithmetic is right, the random numbers are fine, and the p-values are computed by the same formula that works elsewhere.

What the series were

Each one is a random walk: start at zero and add an independent draw from a standard normal at every step.

Xt=Xt−1+εt,εt∼N(0,1) independentX_t = X_{t-1} + \varepsilon_t, \qquad \varepsilon_t \sim N(0, 1) \text{ independent}

Nothing is hidden in there. Two such series built from separate draws are independent by construction, and no amount of staring at them can make one inform the other.

Why the test fails

A random walk has no level to come back to. Whatever value it reaches, the next step starts from there rather than from any home position, so the displacement accumulates instead of averaging out.

That shows up as a variance that grows with time. Measured over 3,000 replications:

time ttvariance of XtX_t
5048.9
200194.1
800757.4

Roughly tt itself, which is the theoretical answer: the sum of tt independent unit-variance steps has variance tt.

Now picture two such series over 200 points. Each wanders somewhere and stays there for a while, because moving away is easy and moving back is not. Over any finite window, two wandering series will very often be drifting in a consistent direction relative to each other, and a scatter plot of one against the other will look like a line.

The regression has no way to discount this. It was built on the assumption that the observations are independent draws, so it counts 200 points of shared drift as 200 independent pieces of evidence. They are closer to one.

One pair would prove nothing: you would be entitled to assume you had been shown the one that happened to line up. So the figure below shows a pair and, beside it, the rejection rate over four hundred more drawn the same way. Press for another pair as often as you like. Then switch to first differences: the same pairs, nothing about the data changed, and the rate falls from four in five to one in twenty.

Interactive: two series with nothing in common

Separate random numbers, no shared input. Press for another pair.

The two series

One against the other

|t| on the slope
4.41
R squared
0.090
At the 5% level
significant
Over 400 pairs
81.8%

|t| is 4.41 against a critical value of 1.972, so this pair is significant - and across 400 independent pairs the test rejects 81.8% of the time, where a valid test on unrelated data should reject 5%. Neither series knows the other exists. The regression counts 200 points of shared drift as 200 independent pieces of evidence, when they are closer to one: a walk has no level to return to, so its variance grows with time and any two walks spend long stretches drifting consistently. Now switch to first differences.

Stationarity, stated

A series is stationary when its statistical behaviour does not depend on when you look at it. The strict version says the joint distribution of any block of observations is unchanged by shifting the block in time. The version used in practice, called weak or covariance stationarity, asks for three things:

  1. the mean E[Xt]=μE[X_t] = \mu does not depend on tt
  2. the variance Var(Xt)=σ2\mathrm{Var}(X_t) = \sigma^2 does not depend on tt
  3. the covariance Cov(Xt,Xt+k)=γk\mathrm{Cov}(X_t, X_{t+k}) = \gamma_k depends on the gap kk only, not on where the window sits

The random walk fails the second and third. Its variance is tt, and its covariance structure moves with the window.

Why this matters more for inference than for description: every estimate borrows strength across observations, and that only makes sense if those observations describe the same process. When the level wanders, an average over the window is an average over different worlds, and a confidence interval around it describes a quantity that never existed.

The repair, and what it costs

A random walk is the running total of its steps, so taking first differences gives the steps back:

∇Xt=Xt−Xt−1=εt\nabla X_t = X_t - X_{t-1} = \varepsilon_t

Those are independent, mean zero, constant variance - exactly the conditions the test needs. Regressing the same 2,000 pairs in first differences gives a significance rate of 4.2%, against the 5% a valid test should produce.

The cost is not zero, and it should be stated rather than buried. The differenced model relates changes in one series to changes in the other. If the question was about levels, it has been replaced by a different question, and the answer should be reported as an answer to that one.

Differencing is also not a universal fix. It removes a stochastic trend, but applied to a series that was already stationary it adds noise for nothing. A straight deterministic trend is better handled by fitting and subtracting it. A variance that grows multiplicatively usually calls for a log first.

The habit this should leave you with

More data does not rescue you here. The spurious rejection rate does not fall as the series lengthens - it climbs, because a longer window gives the two series more room to wander. That is the signature of bias rather than noise, and it is the same shape of problem as confounding: a property of how the data came to exist, not of how much of it there is.

So the first question about any series is whether its behaviour is stable in time, and the honest first tool is a plot. A wandering level or a widening spread is usually visible. Formal tests exist, but they have poor power against slowly moving alternatives, which means a test that fails to reject is weak evidence of stationarity rather than a clearance.

What this sets up

Stationarity is the condition that makes the toolbox meaningful. The next question is what the toolbox then measures: if a series is stationary, its dependence on its own past is a stable, estimable thing, and that dependence is both what makes forecasting possible and what makes ordinary standard errors wrong by a computable factor.

References & further reading

  • Rob J Hyndman, George Athanasopoulos, Forecasting: Principles and Practice, OTexts, 2014source ↗
  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.