What Is Statistical Learning?
The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.
Prerequisites: Probability from Zero
Every supervised model on this site - linear regression, trees, neural networks - is an answer to the same question, posed in the same notation. Before comparing methods it is worth stating that question exactly, because it also tells you which part of your error you can hope to remove and which part you are stuck with no matter how good your method is.
A. The setup
We observe a response and different predictors . We assume there is some relationship between them, written in the general form
Here is a fixed but unknown function representing the systematic information that provides about , and is a random error term, independent of , with mean zero.
Statistical learning is the set of approaches for estimating .
The estimate is written , and the prediction it produces is .
B. Reducible and irreducible error
Suppose for a moment that and are fixed. How wrong is as a prediction of ? Following James et al., the expected squared error decomposes into two conceptually distinct pieces:
The first term is reducible: is not a perfect estimate of , and we can shrink that gap by choosing a better method or gathering more data.
The second term is irreducible, and this is the important one. Even with a perfect estimate - even if exactly - the prediction would still be wrong by , because depends on and is by definition not predictable from .
Why is the irreducible error larger than zero? Two reasons, both worth internalising. First, may contain unmeasured variables that would help predict - since we did not measure them, no built on can use them. Second, it may contain genuinely unmeasurable variation: the same inputs on two different days simply do not produce identical outputs.
The practical consequence is a hard ceiling. is an upper bound on how good any model can get, and it is almost always unknown in practice. A model that appears to beat it is not a triumph; it is a sign you are measuring test error on data the model has already seen.
C. Prediction versus inference
There are two reasons to estimate , and they pull in opposite directions.
For prediction, you only care that is close to . The estimate can be a total black box - you will never inspect it - provided it is accurate.
For inference, you want to understand the relationship: which predictors actually matter, whether each one raises or lowers the response, whether a simple linear summary is adequate. Now cannot be a black box, because the shape of the model is the answer.
This tension recurs constantly:
| Prediction | Inference | |
|---|---|---|
| Goal | Accurate | Understand |
| Model preference | Flexible, possibly opaque | Restrictive, interpretable |
| Typical choice | Ensembles, neural networks | Linear and generalised linear models |
A model chosen purely for accuracy will often be one you cannot explain, and a model chosen for explanation will often leave accuracy on the table. Knowing which of the two you are doing is a prerequisite to choosing sensibly.
D. Parametric and non-parametric estimation
Broadly there are two strategies for producing .
Parametric methods reduce the problem to estimating a fixed set of numbers. You first assume a functional form - most simply, that is linear:
and then only need to estimate the coefficients rather than an arbitrary -dimensional function. That is an enormous simplification. The risk is equally clear: if the true is far from linear, no choice of coefficients will fit it, and the model is wrong in a way more data cannot fix.
Non-parametric methods make no such assumption and let the data determine the shape. They can fit a much wider range of true functions, but because they do not reduce the problem to a few parameters, they need substantially more observations to pin the shape down reliably.
E. Why flexibility is not free
It is tempting to conclude that flexible non-parametric methods are simply better. They are not, for two reasons.
First, as noted, they are hungrier for data. Second - and less obviously - a very flexible method can follow the noise in the training sample as if it were signal. The fit to the data you have looks superb while the performance on data you have not seen degrades.
The figure draws schematic curves, not the polynomial fits below: as you raise Model complexity, training error falls steadily while test error turns back up.
Interactive: the bias-variance tradeoff
Training vs test error as complexity grows.
- Regime
- Well-fit
- Training error
- 0.25
- Test error
- 0.55
Near the sweet spot: the test error is close to its minimum.
Training error always falls as the model grows more flexible, so it is a misleading guide. Test error is bias² + variance + irreducible noise: it bottoms out where the two forces balance, then rises as the model fits noise. Add data (raise the sample size) and the variance term shrinks, pushing the sweet spot to higher complexity and lowering the whole test curve.
The following demonstrates the point on data where the truth is known, so we can compare the fit against itself rather than against noisy observations:
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Running it prints:
degree 1: training MSE 0.1985 error vs true f 0.2465
degree 3: training MSE 0.0542 error vs true f 0.0235
degree 15: training MSE 0.0258 error vs true f 730.5159
Read the two columns against each other. Training MSE falls monotonically with flexibility - degree 15 fits the observed points best of all three, at . Yet its error against the true is : not marginally worse than degree 3's , but roughly thirty thousand times worse.
The blow-up is worth understanding rather than just noting. A degree-15 polynomial fitted to 25 points has enough freedom to weave through the noise, and the wiggles it needs in order to do so become violent near the edges of the interval, where there are fewest points to constrain it. Evaluated on the dense grid, those boundary excursions dominate the average. This is a real and well-known failure mode of high-degree polynomial fits, not an artefact of the seed.
Meanwhile degree 1 shows the opposite failure: its training error is much worse than degree 15's, but it is honest about it - the error against the truth, , is about the same as its training error, because a straight line is too rigid to chase noise. It is simply the wrong shape for a sine wave.
Degree 3 wins on the only column that matters. This gap between "fits the sample" and "captures the truth" is the central phenomenon of the subject, and it has a precise decomposition.
Key takeaways
- Statistical learning estimates an unknown in .
- Prediction error splits into a reducible part, which better methods can shrink, and an irreducible part , which nothing can.
- Irreducible error exists because of unmeasured and unmeasurable variation; it is a hard ceiling on any model's accuracy.
- Prediction favours flexible, opaque models; inference favours restrictive, interpretable ones. Decide which you are doing first.
- Parametric methods assume a form and estimate few parameters; non-parametric methods assume less but need far more data.
- More flexibility always improves training error, and beyond a point worsens accuracy on new data.
What's next
That last observation deserves an exact account rather than an anecdote. The reducible error itself splits into two competing pieces - one that shrinks as models get more flexible and one that grows - and their sum is what you actually pay. That is The Bias–Variance Tradeoff.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.