Skip to content
Kudos AI
Lire en français
Time Series

The Regression That Finds a Relationship That Is Not There

Two series generated from separate random numbers come out significantly related 82.8% of the time, a standard error on dependent data is too narrow by a computable factor of 2.4, and the usual validation split reports a forecaster more than five times better than it is. Three failures, one cause, and the checks that catch each of them.

9 min readKudos AI

Prerequisites: Statistical Inference

Two independently generated random walks drifting together into a convincing scatter plot, the same points in first differences collapsing into a shapeless cloud, and a validation split sliding from shuffled to forward as the error bar jumps more than five-fold.

Write ten lines of code. Generate one series of 200 numbers by starting at zero and adding a random step each time. Generate a second series the same way, from separate random draws, so that by construction neither knows anything about the other. Regress one on the other and test the slope at the usual 5% level.

Do that 2,000 times and the slope comes out significant in 82.8% of them.

A valid test on unrelated data should reject 5% of the time. This one rejects sixteen times too often, and every p-value is computed by the formula that works everywhere else.

That result is the entrance to a subject, and the subject is what happens to ordinary statistics when the rows are in an order that matters.

Failure one: the level that never comes back

Each of those series is a random walk: Xt=Xt−1+εtX_t = X_{t-1} + \varepsilon_t with independent standard-normal steps. The key property is that it has no level to return to. Whatever value it reaches, the next step starts from there, so displacement accumulates instead of averaging away.

That shows up as a variance which grows with time. Measured over 3,000 replications:

time ttvariance of XtX_t
5048.9
200194.1
800757.4

Roughly tt itself, which is the theoretical answer for a sum of tt independent unit-variance steps.

Now take two such series over any finite window. Each wanders somewhere and stays a while, because leaving is easy and returning is not, so the two will very often be drifting in a consistent direction relative to each other. A scatter of one against the other looks like a line. The regression, built on the assumption of independent draws, counts 200 points of shared drift as 200 independent pieces of evidence when they are closer to one.

The condition being violated is stationarity: the requirement that the statistical behaviour of a series does not depend on when you look at it. In its working form it asks for a mean that does not depend on tt, a variance that does not depend on tt, and a covariance Cov(Xt,Xt+k)\mathrm{Cov}(X_t, X_{t+k}) that depends on the gap kk only. A random walk fails the last two.

Why that matters for inference more than for description: every estimate borrows strength across observations, which only makes sense if those observations describe the same process. When the level wanders, an average over the window is an average over different worlds, and an interval around it describes a quantity that never existed.

The standard repair is first differencing. The differences of a random walk are exactly the independent steps it was built from, so the assumption is restored and the same 2,000 pairs now reject at 4.2%, which is what a valid 5% test should produce. The cost should be stated rather than buried: the model now relates changes to changes, which is a different question from the one about levels.

The uncomfortable part is what more data does. The spurious rejection rate does not fall as the series lengthens. It rises, because a longer window gives both series more room to wander. That is the signature of bias rather than noise, and it makes this the same shape of problem as confounding: a property of how the data came to exist, not of how much of it there is.

Failure two: the standard error that is too narrow

Suppose you do have a stationary series. Dependence has not gone away, and it now has a price you can compute exactly.

The simplest process with memory is AR(1): Xt=φXt−1+εtX_t = \varphi X_{t-1} + \varepsilon_t with ∣φ∣<1|\varphi| < 1. Multiply through by Xt−kX_{t-k} and take expectations, and the noise term drops out, leaving autocorrelations that decay geometrically:

ρk=φ k\rho_k = \varphi^{\,k}

Measured on 39,800 simulated points with φ=0.7\varphi = 0.7, the autocorrelations at lags 1, 2, 3 and 5 come out as 0.7039, 0.4980, 0.3527 and 0.1796, against a theory of 0.7, 0.49, 0.343 and 0.1681. Agreement to within 0.012.

Now the consequence. The variance of the sample mean of such a series is not σ2/n\sigma^2/n. Summing covariances over all pairs gives an inflation factor that, for the AR(1), collapses to a geometric series:

1+2∑k=1∞φ k=1+φ1−φ1 + 2\sum_{k=1}^{\infty}\varphi^{\,k} = \frac{1+\varphi}{1-\varphi}

At φ=0.7\varphi = 0.7 that is 5.66675.6667. Two readings, both worth carrying:

  • the standard error is a square root, so computing it as though the points were independent understates it by 5.6667≈2.4\sqrt{5.6667} \approx 2.4
  • the effective sample size is nn divided by 5.6667, so 6,000 dependent points carry about as much information about the mean as 1,060 independent ones

Nothing in the output announces this. The estimate is not even biased. It is a right answer reported with far more confidence than it earned, which is why this failure is more common than the dramatic one: it never looks wrong.

The direction checks out at the edges. At φ=0\varphi = 0 the factor is 1, the independent case. As φ→1\varphi \to 1 it diverges, which is the random walk having no usable mean at all. Negative φ\varphi gives a factor below 1, because alternating errors partly cancel and such a series carries more information per point than an independent one.

One noise sequence drives the figure below, and the dial only changes how long the series remembers it. Watch two things at once. The autocorrelation bars sit on the theoretical curve, which is what makes the plot a diagnostic rather than a decoration. And the effective sample size falls away while the series above it goes on looking like a series - no warning appears on the surface of the output, which is exactly why this failure is the common one.

Interactive: the same shocks, a longer memory

One noise sequence throughout. Only how long it is remembered changes.

The series

Autocorrelation: measured against φᵏ

-101012
measuredtheory φᵏ
Measured ρ₁
0.6970
Variance inflation
5.6667
Effective n, from 6000
1059
Interval too narrow by
2.38x

The measured bars sit on the theoretical φᵏ curve, which is what makes the autocorrelation function a diagnostic rather than a decoration: a geometric decay is the fingerprint of an AR(1), and its rate is φ itself. The variance of the sample mean is inflated by 5.6667, so 6,000 dependent points are worth about 1059 independent ones. Push φ up and watch that number fall while the picture above barely changes.

Failure three: the split that measures the wrong task

Cross-validation is the most reliable habit in applied machine learning, and on a time series it can report a model more than five times better than it is.

Take a dependent series of 240 points, hold out 48, and measure the simplest method for each task - fill a held-out point in from the training points around it, or carry the last value forward:

splitRMSE
48 points held out at random0.7515
the last 48, trained on the first 1924.2074

Random splitting is correct when rows are exchangeable - when a row's position carries no information. On a series with memory, position is most of the information. Shuffle, and every held-out point lands between training points, two in three of them with one right beside it on each side, so the model is no longer being asked what happens next. It is being asked to fill a gap between two known values, which on a dependent series is easy.

The 0.75 is therefore not a mildly optimistic estimate of forecasting skill. It is an accurate answer to a question nobody asked.

The rule that prevents this is simple to state - no information from after the forecast origin may reach the model - and it catches more than shuffling. Scaling computed over the whole series lets the training data know the future mean and variance. Imputing a gap from later observations smuggles interpolation into the training set. A centred rolling mean includes the point it is meant to predict. And choosing a model order by looking at the test period spends the test set without recording that it was spent, which is the easiest of the four to commit because it happens across sessions rather than inside one script.

The figure below runs both hold-outs on its own series. It is an independent implementation with an independent generator, and it reads 0.7557 and 4.2224 against the 0.7515 and 4.2074 above, which is the useful thing to notice: the factor of 5.6 belongs to the procedure, not to anyone's seed. Widen the hold-out and watch the two diverge further, because forecast error accumulates with the horizon and gap-filling barely grows.

Interactive: one series under two hold-outs

One series with memory. Only the split changes.

Shuffled hold-out
0.7557
Forecast forward
4.2224
Flattered by
5.6x
Exact, pooled
4.9497

Holding out 48 points at random gives an RMSE of 0.7557; holding out the last 48 gives 4.2224. A factor of 5.6, on the same data. Shuffle a series with memory and every test point sits between training points, two in three with one right beside it on each side, so the model is no longer asked what happens next - it is asked to fill a gap, which is easy here. Drag the hold-out wider: the gap-filling error rises a little as the gaps lengthen, while the forecast error grows far faster, because forecast error accumulates with the horizon.

The baseline that settles arguments

One more measurement, because it costs a line and prevents a category of self-deception. Forecasting 40 points from 120 of history on a random walk:

forecastRMSE
the last observed value4.0419
the mean of the training window6.5258
extrapolating the fitted drift4.4063

The simplest wins. The mean assumes the series returns to a level it does not have. The drift extrapolates a slope that is just the accumulated noise of the training window divided by its length - real in the sample, absent from the process.

So put the naive forecast in every comparison. If a method cannot beat "tomorrow looks like today", its apparent skill is coming from somewhere other than the data, and very often from a split like the one above.

What to actually do

Four habits cover most of it.

  1. Plot the series before anything else. A wandering level or a widening spread is usually visible. Formal stationarity tests have poor power against slowly moving alternatives, so a test that fails to reject is weak evidence rather than a clearance.
  2. Difference when the level wanders, and say so. Report that the model describes changes. Do not difference a series that was already stationary - that adds noise for nothing.
  3. Never report a standard error from dependent data without correcting it. For anything close to AR(1), (1+φ)/(1−φ)(1+\varphi)/(1-\varphi) is the factor, and the effective sample size is the honest number to quote.
  4. Split forward, at several origins. Every training set ends before its test set begins, and repeating at several cut points shows how performance varies with when the forecast was made. Report the spread, not only the average.

Three failures, one cause. None of them announces itself in the output, and none is fixed by collecting more data. What fixes them is knowing the rows are ordered, and letting that order constrain the test you run, the interval you report, and above all the split you evaluate on.

References & further reading

  • Rob J Hyndman, George Athanasopoulos, Forecasting: Principles and Practice, OTexts, 2014source ↗
  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

4 min readTime Series

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

StatisticsMachine Learning
6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
7 min readAnomaly Detection

The Detector That Never Fires Is 99.5% Accurate

At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

Machine LearningStatistics
← Back to all articles