Skip to content
Kudos AI

Cross-Validation

A resampling method that estimates a model’s test error by repeatedly fitting it on part of the data and evaluating it on the part held out.

Also known as: k-fold cross-validation

Understanding Cross-Validation

Training error systematically understates future error, because the model was tuned to the very observations it is being scored on. The honest alternative is to evaluate on data withheld from fitting. A single split does this but wastes data and is noisy: the estimate depends on which observations happened to fall into the held-out portion.

k-fold cross-validation resolves both problems. The data is partitioned into k roughly equal folds. The model is fitted k times; on each round one fold is held out for validation and the remaining k − 1 are used for training. Averaging the k validation scores gives the cross-validation estimate. Every observation is used for validation exactly once and for training k − 1 times, so no data is wasted.

The choice of k involves its own trade-off. With k equal to the number of observations, known as leave-one-out cross-validation, each model is fitted on almost the entire dataset, so the estimate has low bias, but the k fitted models are highly correlated with one another, which inflates the variance of the average, and the computational cost is high. Smaller k reduces both correlation and cost at the price of some bias. Values of 5 and 10 are the conventional compromise.

The most common practical mistake is leakage. If feature scaling, feature selection, or imputation is fitted on the full dataset before splitting, information from the validation fold has already influenced the model, and the resulting score is optimistic. Every data-dependent step must be re-fitted within each training fold.

How to Calculate

CV(k) = (1/k) Σᵢ₌₁ᵏ Errᵢ

where

k
the number of folds
Errᵢ
the error measured on fold i, from a model trained on the other k − 1 folds
CV(k)
the cross-validated estimate of test error

Example of Cross-Validation

With 100 observations and k = 5, the data is divided into five folds of 20. The model is fitted five times, each on 80 observations and scored on the remaining 20, and the five scores are averaged.

The procedure is most valuable for choosing hyperparameters. To select a regularization strength, run the whole five-fold procedure once per candidate value and pick the value with the best average score. Because each candidate is evaluated on data it was not fitted on, the comparison is fair in a way that training error could never be.

A useful refinement is the one-standard-error rule: rather than taking the single best mean score, choose the simplest model whose score lies within one standard error of the best. Differences smaller than the noise in the estimate are not real evidence, and the simpler model usually carries less variance.

Advantages and Disadvantages

Pros

  • Uses every observation for both training and validation, which matters most when data is scarce.
  • Produces a far more stable estimate than a single train/test split.
  • Provides a spread across folds, indicating how uncertain the estimate itself is.

Cons

  • Costs k model fits instead of one, which can be prohibitive for large models.
  • Assumes observations are exchangeable, so plain k-fold is wrong for time series or grouped data.
  • Easy to invalidate through leakage if preprocessing is done outside the fold loop.

Frequently Asked Questions

Why is k = 5 or k = 10 recommended rather than leave-one-out?

Leave-one-out fits on nearly all the data, so its estimate has low bias, but the fitted models are almost identical to one another, which makes the averaged estimate high-variance, and it requires as many fits as there are observations. Values of 5 and 10 have been found empirically to suffer from neither excessive bias nor excessive variance.

Can cross-validation be used on time-series data?

Not in its standard form. Random folds would train on the future to predict the past. Time series require forward-chaining schemes in which the validation period always follows the training period.

After cross-validating, which of the k models should be deployed?

None of them individually. Cross-validation is a procedure for estimating error and selecting settings. Once the settings are chosen, the final model is re-fitted on the entire dataset.

The Bottom Line

Cross-validation buys a reliable estimate of out-of-sample performance at the cost of repeated fitting. It is the standard tool for model selection precisely because it never scores a model on data that model was fitted on, provided every data-dependent step is kept inside the fold.