The Bias-Variance Tradeoff
The exact decomposition of expected test error into squared bias, variance, and irreducible noise, demonstrated numerically with a 2,000-run simulation where all three terms are measured separately and shown to add up.
Prerequisites: What Is Statistical Learning?
What Is Statistical Learning? showed a degree-15 polynomial fitting its training sample better than any competitor while being catastrophically wrong about the underlying function. That was a demonstration. This article gives the account: the expected test error splits into exactly three pieces, two of which we control and one of which we do not, and the two we control move in opposite directions.
A. The decomposition
Take a test point , and let be a model fitted to a random training sample. The expected squared error at , averaged over all the training samples we might have drawn, decomposes as
where the bias is .
This is an equality, not an approximation or a heuristic. Each term means something concrete:
- Bias is the error from approximating a complicated real problem with a simpler model. A straight line fitted to a curve is biased everywhere, and no amount of extra data removes that.
- Variance is how much would change if you refitted it on a different training sample of the same size. A method that swings wildly from sample to sample has high variance, and any single fit is unreliable.
- Irreducible error is , the floor from the previous article.
Since squared bias and variance are both non-negative, their sum is non-negative, so expected test error can never fall below . The floor is real.
B. Why the two terms fight
The general pattern: as a method becomes more flexible, bias falls and variance rises.
More flexibility means the model can bend towards the true , so it is less systematically wrong - bias down. But it also means the model can bend towards the particular noise in this sample, so a different sample would produce a noticeably different fit - variance up.
Test error is the sum. Early on, flexibility buys a large bias reduction for a small variance increase and total error falls. Past some point the trade reverses and total error climbs. The best model sits where the two rates of change cancel.
The figure draws schematic curves in arbitrary units, not the simulation below; its noise floor sits at 0.16, not the 0.09 used there.
Interactive: the bias-variance tradeoff
Training vs test error as complexity grows.
- Regime
- Well-fit
- Training error
- 0.25
- Test error
- 0.55
Near the sweet spot: the test error is close to its minimum.
Training error always falls as the model grows more flexible, so it is a misleading guide. Test error is bias² + variance + irreducible noise: it bottoms out where the two forces balance, then rises as the model fits noise. Add data (raise the sample size) and the variance term shrinks, pushing the sweet spot to higher complexity and lowering the whole test curve.
C. Measuring all three terms
The decomposition is usually presented as theory because in practice you cannot compute it - you do not know and you have only one training sample. But in a simulation we control both, so we can measure each term separately and check they really do add up.
The setup: , noise so , training samples of 30 points, and a single test point . We refit 2,000 times on fresh samples and watch what the predictions at do.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Running it prints:
true f(x0) = 0.8090, Var(eps) = 0.0900
degree 1: bias^2 0.2684 variance 0.0130 + noise 0.0900 = 0.3713
degree 3: bias^2 0.0059 variance 0.0105 + noise 0.0900 = 0.1064
degree 9: bias^2 0.0000 variance 0.0295 + noise 0.0900 = 0.1195
D. Reading the table
Every claim in section B is visible here as a number.
Bias collapses as flexibility grows. Degree 1 has squared bias - a straight line simply cannot pass near at , and its average prediction over 2,000 fits was against a true value of . By degree 3 the squared bias is , and by degree 9 it rounds to zero: on average, the flexible model is exactly right.
Variance grows as flexibility grows. It more than doubles from at degree 1 to at degree 9. The degree-9 fit is right on average but any individual fit is noticeably off, in a direction that depends on which 30 points it happened to see.
The sum is minimised in the middle. Total error runs . Degree 3 wins - not because it is best on either term individually (degree 9 has lower bias) but because it is the best compromise.
The floor holds. No configuration gets below . The best total, , is only above the irreducible noise, and that remaining gap is the reducible error still on the table.
A caution about the degree-9 row. Squared bias printing as does not mean the model is unbiased everywhere - only at this particular , and only to four decimal places after averaging 2,000 fits. Bias is a function of ; at a point near the boundary of the interval the same model would show substantial bias, for exactly the reason the degree-15 polynomial exploded in the previous article.
E. Does the decomposition actually hold?
The three terms should sum to the expected test error, which we can measure directly by generating fresh noisy observations at and averaging the squared prediction errors. Doing that alongside the decomposition gives:
| degree | bias² + var + noise | measured test MSE |
|---|---|---|
| 1 | 0.3713 | 0.3764 |
| 3 | 0.1064 | 0.1072 |
| 9 | 0.1195 | 0.1171 |
The two columns agree to within a few thousandths - the residual difference is Monte Carlo error from using 2,000 simulations rather than infinitely many. The identity is exact; our estimate of it is merely very good.
F. What this means in practice
You cannot compute bias and variance on real data, because you have one sample and no access to . What you can do is recognise their signatures:
| Symptom | Likely cause | Response |
|---|---|---|
| High training error, similar test error | High bias (underfitting) | More flexible model, better features |
| Very low training error, much higher test error | High variance (overfitting) | Regularise, simplify, more data |
| Both errors near the noise floor | Near the optimum | Stop |
Note the asymmetry in the fixes: more data reduces variance but does nothing for bias. Doubling the sample will not help a straight line fit a sine wave. If your model is biased, you need a different model, not a bigger dataset.
Key takeaways
- Expected test error decomposes exactly into squared bias, variance, and irreducible noise.
- Bias falls and variance rises with flexibility; the sum is U-shaped and the best model is the compromise, not the extreme.
- In simulation the terms are measurable: degree 3 won with total against (biased) and (high-variance).
- Test error can never go below .
- More data cures variance, not bias. Diagnose which you have before choosing a remedy.
What's next
The diagnostic table above needs an honest estimate of test error, and the training error emphatically is not one. Getting that estimate from the data you already have - without a separate test set you may not be able to afford - is the job of Cross-Validation and Resampling.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.