Skip to content
Kudos AI
Lire en français
Time Series

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

4 min readKudos AI

Prerequisites: The Regression That Finds a Relationship That Is Not There

A series split at random, with test points sitting between their own neighbours in the training set, and then split forward in time so every test point lies beyond the last thing the model saw.

Take 600 steps of a random walk: yt=yt−1+εty_t = y_{t-1} + \varepsilon_t with independent standard normal steps. The increments are independent by construction, so there is nothing in this series to forecast. The best possible prediction of yt+1y_{t+1}, knowing everything up to tt, is yty_t.

Fit a five-nearest-neighbour regression of yy on the time index. Evaluate it with random five-fold cross-validation, which is the default in every library.

R2=0.9983,RMSE=0.865.R^2 = 0.9983, \qquad \text{RMSE} = 0.865.

A. What the random split measured

Nothing about forecasting. Under a random split, the neighbours of a held-out point in time are overwhelmingly in the training set: point tt is being predicted from t−2t-2, t−1t-1, t+1t+1, t+2t+2. The model is not extrapolating, it is interpolating between two observations it was shown, and on a continuous series that is close to reading the answer off the page.

Evaluate the same model forward in time instead. Train on everything up to a cut, predict the block after it, move the cut, repeat:

EvaluationR2R^2RMSE
random five-fold0.99830.865
forward chaining0.655910.760
last value carried forward0.688110.244

The RMSE is 12.44 times larger under the honest split. And the third row is the one that should end the discussion: simply repeating the last observed value scores better than the model, on both measures. The model has negative value, and the first evaluation reported it as nearly perfect.

The figure runs a simpler version of the same trap on 240-step walks: under the shuffled split each held-out point is filled in from its two neighbours, against carrying the last value forward under the honest one, so the flattering factor is about 6 rather than the 12.44 measured here.

Interactive: one series under two hold-outs

One series with memory. Only the split changes.

Shuffled hold-out
0.7557
Forecast forward
4.2224
Flattered by
5.6x
Exact, pooled
4.9497

Holding out 48 points at random gives an RMSE of 0.7557; holding out the last 48 gives 4.2224. A factor of 5.6, on the same data. Shuffle a series with memory and every test point sits between training points, two in three with one right beside it on each side, so the model is no longer asked what happens next - it is asked to fill a gap, which is easy here. Drag the hold-out wider: the gap-filling error rises a little as the gaps lengthen, while the forecast error grows far faster, because forecast error accumulates with the horizon.

B. Why R2=0.6559R^2 = 0.6559 is not a consolation

The forward-chaining R2R^2 still looks respectable, and it should not be read as "the model captures two thirds of the variation". A random walk wanders, so most of the variance in any test block is the level it drifted to, and predicting the level approximately right is worth a large R2R^2 while being worth nothing at all.

This is why forecasting is scored against a benchmark rather than against a constant mean. Against "carry the last value forward" the model is behind, which is the correct reading: there was no signal, the model found none, and every number that suggested otherwise came from how it was measured.

C. Where this happens outside a demonstration

The random walk makes the effect impossible to argue with, but the mechanism needs only autocorrelation, which every real series has.

  • Random splits of any temporal data. Sensor readings, prices, demand, server load, patient vitals. Neighbouring rows are similar, so a random split puts near-duplicates of each test row in the training set.
  • Features built over the whole series. A centred rolling mean, a standardisation using the full-sample mean and standard deviation, a seasonal decomposition, an imputation fitted on everything. Each one moves information backwards in time, and each survives a random split untouched.
  • Grouped data split at the row level. Multiple rows per customer, per patient, per session. The split has to be by group, and time is the same argument with the group being "the same week".
  • Tuning on the test block. Forward chaining is honest about the split and can still be spent, one hyperparameter at a time, on the last block.

D. The checks that are actually cheap

  • Split forward, always, for anything with a time index. It costs one line. If the model is genuinely time-invariant the two splits will agree, and the agreement is itself worth knowing.
  • Always report a naive benchmark. Last value for a level, seasonal naive for a seasonal series. A model that does not beat it is not a model yet, and the comparison costs nothing.
  • Recompute every feature inside the fold. If a transformation touches rows that the model is not allowed to have seen, it belongs inside the training step, not in the data-preparation notebook.
  • Ask what the score would be if the series were unpredictable. On this data, 0.9983. If your evaluation can return that number on a random walk, it is not measuring what you think.

References & further reading

  • Rob J Hyndman, George Athanasopoulos, Forecasting: Principles and Practice, OTexts, 2014source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

9 min readTime Series

The Regression That Finds a Relationship That Is Not There

Two series generated from separate random numbers come out significantly related 82.8% of the time, a standard error on dependent data is too narrow by a computable factor of 2.4, and the usual validation split reports a forecaster more than five times better than it is. Three failures, one cause, and the checks that catch each of them.

StatisticsMachine Learning
7 min readStatistical Learning Foundations

Cross-Validation and Resampling

Why training error is a biased estimate of test error, and how the validation set, leave-one-out, and k-fold approaches fix it, with a five-fold LOOCV computation worked out observation by observation.

StatisticsMachine Learning
6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
← Back to all articles