A Score That Loses to Doing Nothing
A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.
Prerequisites: The Regression That Finds a Relationship That Is Not There
Take 600 steps of a random walk: with independent standard normal steps. The increments are independent by construction, so there is nothing in this series to forecast. The best possible prediction of , knowing everything up to , is .
Fit a five-nearest-neighbour regression of on the time index. Evaluate it with random five-fold cross-validation, which is the default in every library.
A. What the random split measured
Nothing about forecasting. Under a random split, the neighbours of a held-out point in time are overwhelmingly in the training set: point is being predicted from , , , . The model is not extrapolating, it is interpolating between two observations it was shown, and on a continuous series that is close to reading the answer off the page.
Evaluate the same model forward in time instead. Train on everything up to a cut, predict the block after it, move the cut, repeat:
| Evaluation | RMSE | |
|---|---|---|
| random five-fold | 0.9983 | 0.865 |
| forward chaining | 0.6559 | 10.760 |
| last value carried forward | 0.6881 | 10.244 |
The RMSE is 12.44 times larger under the honest split. And the third row is the one that should end the discussion: simply repeating the last observed value scores better than the model, on both measures. The model has negative value, and the first evaluation reported it as nearly perfect.
The figure runs a simpler version of the same trap on 240-step walks: under the shuffled split each held-out point is filled in from its two neighbours, against carrying the last value forward under the honest one, so the flattering factor is about 6 rather than the 12.44 measured here.
Interactive: one series under two hold-outs
One series with memory. Only the split changes.
- Shuffled hold-out
- 0.7557
- Forecast forward
- 4.2224
- Flattered by
- 5.6x
- Exact, pooled
- 4.9497
Holding out 48 points at random gives an RMSE of 0.7557; holding out the last 48 gives 4.2224. A factor of 5.6, on the same data. Shuffle a series with memory and every test point sits between training points, two in three with one right beside it on each side, so the model is no longer asked what happens next - it is asked to fill a gap, which is easy here. Drag the hold-out wider: the gap-filling error rises a little as the gaps lengthen, while the forecast error grows far faster, because forecast error accumulates with the horizon.
B. Why is not a consolation
The forward-chaining still looks respectable, and it should not be read as "the model captures two thirds of the variation". A random walk wanders, so most of the variance in any test block is the level it drifted to, and predicting the level approximately right is worth a large while being worth nothing at all.
This is why forecasting is scored against a benchmark rather than against a constant mean. Against "carry the last value forward" the model is behind, which is the correct reading: there was no signal, the model found none, and every number that suggested otherwise came from how it was measured.
C. Where this happens outside a demonstration
The random walk makes the effect impossible to argue with, but the mechanism needs only autocorrelation, which every real series has.
- Random splits of any temporal data. Sensor readings, prices, demand, server load, patient vitals. Neighbouring rows are similar, so a random split puts near-duplicates of each test row in the training set.
- Features built over the whole series. A centred rolling mean, a standardisation using the full-sample mean and standard deviation, a seasonal decomposition, an imputation fitted on everything. Each one moves information backwards in time, and each survives a random split untouched.
- Grouped data split at the row level. Multiple rows per customer, per patient, per session. The split has to be by group, and time is the same argument with the group being "the same week".
- Tuning on the test block. Forward chaining is honest about the split and can still be spent, one hyperparameter at a time, on the last block.
D. The checks that are actually cheap
- Split forward, always, for anything with a time index. It costs one line. If the model is genuinely time-invariant the two splits will agree, and the agreement is itself worth knowing.
- Always report a naive benchmark. Last value for a level, seasonal naive for a seasonal series. A model that does not beat it is not a model yet, and the comparison costs nothing.
- Recompute every feature inside the fold. If a transformation touches rows that the model is not allowed to have seen, it belongs inside the training step, not in the data-preparation notebook.
- Ask what the score would be if the series were unpredictable. On this data, 0.9983. If your evaluation can return that number on a random walk, it is not measuring what you think.
References & further reading
- Rob J Hyndman, George Athanasopoulos, Forecasting: Principles and Practice, OTexts, 2014source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.