Understanding Bias-Variance Trade-off
For a squared-error loss, the expected test error at a point decomposes into three additive pieces, and the decomposition is an identity rather than an approximation. The first is the square of the bias: how far the model’s average prediction, across all the training sets it might have been fitted on, sits from the truth. The second is the variance: how much the prediction moves as the training set changes. The third is irreducible noise inherent in the data.
Bias measures systematic error introduced by the model’s assumptions. A linear model applied to a genuinely non-linear relationship is biased no matter how much data it is given, because no straight line can represent a curve. Variance measures instability. A model with enough flexibility to chase individual observations will produce a noticeably different fit on a different sample from the same population.
The two move in opposite directions as flexibility increases. Adding capacity lets a model track the true function more closely, lowering bias, while simultaneously letting it track the noise, raising variance. Test error therefore traces a U-shape: it falls while the bias reduction dominates, reaches a minimum, then rises as the variance increase takes over. Training error, by contrast, falls monotonically, which is exactly why it cannot be used to select complexity.
The irreducible term sets a floor under the expected error. Even a model that recovered the true function perfectly would still make errors, because the observations themselves contain randomness. A measured error on a finite test set scatters either side of that floor, so only a reported accuracy that beats the floor by more than that sampling noise is a sign of a leak between training and evaluation data rather than of a superior model.
How to Calculate
E[(y − f̂(x))²] = Bias[f̂(x)]² + Var[f̂(x)] + σ²
where
- f̂(x)
- the model’s prediction at x, itself random because the training set is random
- Bias[f̂(x)]
- E[f̂(x)] − f(x): the gap between the average prediction and the truth
- Var[f̂(x)]
- the variability of the prediction across possible training sets
- σ²
- the irreducible variance of the noise in y
Example of Bias-Variance Trade-off
Contrast two extreme models on the same problem. A model that ignores its input entirely and always predicts a fixed constant has zero variance, since a new training set changes nothing, but very large bias, since it cannot respond to x at all.
At the other extreme, a model that interpolates every training point exactly has near-zero bias on average but enormous variance: re-draw the training sample and the fitted function changes shape completely, because it is following the noise.
Neither is useful. The model that minimizes test error sits between them, and its location depends on how much data is available: as the sample grows, the variance cost of extra flexibility falls, so the optimum shifts toward more complex models. This is why a model class that overfits badly on a thousand examples can be exactly right on a million.
Frequently Asked Questions
Can a model have both high bias and high variance?
Yes. A badly specified model can be simultaneously wrong on average and unstable across samples. It usually indicates that the model class is inappropriate for the problem, rather than that a hyperparameter needs adjusting.
Why can I not just compute the bias and variance for my model?
Both terms are defined with respect to the true underlying function and the distribution over possible training sets, neither of which is observable in practice. This is why the trade-off is navigated indirectly, by estimating test error through cross-validation or a held-out set.
Does the trade-off mean that flexible models are dangerous?
It means flexibility has a cost that must be paid for with data or constrained by regularization. Given enough data, or an effective penalty on complexity, highly flexible models can achieve both low bias and controlled variance.
The Bottom Line
Every model’s expected error is squared bias plus variance plus irreducible noise. Because flexibility reduces the first while inflating the second, model selection is the search for the point where their sum is smallest, and that point can only be located using data held out from fitting.