Understanding Stationarity
Statistical procedures borrow strength across observations, which only makes sense if those observations describe the same process. Stationarity is that condition made precise: the distribution generating the series does not shift as time passes. When it does shift, an average across the window is an average over different worlds, and an interval around it describes a quantity that never existed.
The clearest demonstration is the spurious regression. Generate two random walks from entirely separate noise, so that by construction neither knows anything about the other, and regress one on the other. Across 2,000 such pairs of 200 points, the slope is significant at the 5% level 82.8% of the time. Nothing is wrong with the arithmetic and nothing is wrong with the random numbers; the test is being asked a question it cannot answer.
The mechanism is the absence of a level to return to. A random walk is the running total of independent steps, so its variance grows with time: 48.9 at t = 50, 194.1 at t = 200, 757.4 at t = 800. Over any finite window, two such series wander, and any two wandering series will drift together for long stretches. A test built for independent observations reads that shared drift as a relationship.
Differencing repairs it because the differences of a random walk are exactly the independent steps it was built from. Regressing the same pairs in first differences gives a significance rate of 4.2%, which is what a valid 5% test should produce. The cost is that the model now describes changes rather than levels, which is a different question and should be reported as one.
How to Calculate
strict: (X_t, …, X_{t+k}) ~ (X_{t+h}, …, X_{t+k+h}) for all h; weak: E[X_t] = μ, Var(X_t) = σ², Cov(X_t, X_{t+k}) = γ_k
where
- μ, σ²
- a mean and variance that do not depend on t
- γ_k
- a covariance depending only on the gap k, not on when the window sits
- unit root
- the non-stationary case where shocks accumulate instead of fading
- ∇X_t = X_t − X_{t−1}
- first differencing, which turns a random walk back into its steps
Example of Stationarity
Two independent random walks of 200 points, regressed on each other 2,000 times: the slope is significant at the 5% level in 82.8% of the trials. The same pairs in first differences: 4.2%.
The variance of a random walk at time t, measured over 3,000 replications: 48.9 at t = 50, 194.1 at t = 200 and 757.4 at t = 800 - growing in proportion to t rather than settling at any value.
On a random walk, forecasting the last observed value gives RMSE 4.0419 against 6.5258 for the training mean, because there is no mean to revert to for the mean forecast to exploit.
Frequently Asked Questions
How do I tell whether a series is stationary?
Plot it first - a wandering level or a widening spread is usually visible. Formal tests exist, but they have low power against slowly-moving alternatives, so a test that fails to reject is weak evidence of stationarity rather than a clearance.
Is differencing always the answer?
No. It fixes a stochastic trend and over-differences a series that was already stationary, adding noise. A deterministic trend is better removed by fitting it, and a changing variance often calls for a transformation instead.
Does more data help with a non-stationary series?
Not for this problem. The spurious rejection rate does not fall as the series lengthens - it rises, because the two series have more room to wander. This is bias rather than noise, which is what makes it dangerous.
The Bottom Line
Stationarity is the assumption that makes the rest of the toolbox meaningful, and the cost of ignoring it is not a wider interval but a confident answer to a question nobody asked. Check it before anything else, and treat the repair as a change of question rather than a formality.