Basis Functions and Piecewise Polynomials
One idea covers polynomial and step-function regression, and extends to anything else you can write down: transform the predictor, then fit a linear model in the transformed columns.
Linear regression assumes the effect of a predictor is a straight line. When it is not, the usual instinct is to reach for a different algorithm. There is a cheaper move: keep least squares and change what you regress on.
The basis function idea
Choose a family of transformations of , fixed and known in advance:
and fit
Two familiar methods are instances of this.
- Polynomial regression takes .
- Step functions cut the range at and take , an indicator that is 1 inside the bin and 0 outside. This is a piecewise constant fit, popular where natural breakpoints exist - five-year age bands, for instance - and poor where they do not, since a bin cannot show a trend inside itself.
The crucial word is fixed. The are chosen before fitting, not estimated, so the model is a linear model in the transformed columns. Least squares applies unchanged, and so does everything built on it: standard errors, -statistics, confidence intervals. You have bought curvature without leaving the linear model.
Piecewise polynomials
Instead of one global polynomial, fit a separate low-degree polynomial in each region. Place knots through the range and fit cubics, one per region. With one knot at :
Each piece is an ordinary cubic fit on its own subset. Two cubics, four parameters each: eight degrees of freedom.
What goes wrong
Nothing forces the two pieces to agree at the knot, so they do not. Fitting this on 200 simulated points with knots at 40, 50 and 60 - four cubics, sixteen parameters - the fitted value jumps at every knot, each jump taken before rounding:
These are not rounding artefacts. The curve is genuinely discontinuous, and as a description of a smooth relationship it is nonsense.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Constraints, and what they cost
The fix is not fewer knots but constraints. Require the pieces to meet:
- Continuity at the knot. The V-shaped join this produces still looks wrong.
- Continuous first derivative - no corner.
- Continuous second derivative - no visible change in curvature.
Each constraint removes one free parameter. Starting from eight with one knot:
The result is a cubic spline. In general a cubic spline with knots uses
degrees of freedom, because each additional knot adds a cubic (four parameters) and three constraints, a net cost of one.
The general definition. A degree- spline is a piecewise degree- polynomial with continuous derivatives up to order at each knot. A linear spline is continuous with a corner allowed; the piecewise constant functions above are degree-0 splines, where even continuity is dropped.
Cubic is the usual choice for a simple reason: a discontinuity in the third derivative is not something the eye can detect.
Here is that accounting as a dial, run on the two hundred points above. At zero constraints the four cubics jump by the three numbers in the table; each constraint then removes one parameter per knot, walking the count 16, 13, 10, 7 and ending at the K + 4 of a cubic spline.
Two things the figure shows that the argument above leaves out. Those nine parameters buy almost nothing: the residual sum of squares rises by under three per cent across the whole journey, so the constraints are close to free. And the corner gets worse before it gets better. Imposing continuity alone more than doubles the sharpest change of slope, from 2.42 to 5.95, which is the V-shaped join mentioned above, measured: making the pieces touch is not the same as making them agree which way to go.
Interactive: the join, and what it costs to close it
Each constraint removes exactly one parameter.
- Degrees of freedom
- 16
- Largest jump
- 6.34
- Sharpest corner
- 2.42
- Residual sum of squares
- 11232.55
With nothing imposed the four cubics jump 6.34, 5.55 and 6.26 at the three knots, which is the lesson’s table. Each constraint then removes exactly one parameter per knot, so the count walks 16, 13, 10, 7 and ends at the K + 4 of a cubic spline. Right now the fit has 16 degrees of freedom, a largest jump of 6.34 and a sharpest corner of 2.42. Two things the numbers say and the prose does not. Those nine parameters are nearly free: the residual sum of squares rises only from 11232.55 to 11534.94, under three per cent, so almost nothing that was bought with them was worth having. And the corner gets worse before it gets better, 2.42 becoming 5.95 at one constraint, because making the pieces touch is not the same as making them agree which way to go.
Where this is heading
We now have a fitting problem with constraints, which is less convenient than plain least squares. The next lesson removes that inconvenience entirely: one well-chosen extra column per knot enforces all three constraints automatically, turning the whole thing back into an ordinary regression.
Before the quiz
Be able to write polynomial and step regression as basis models, say why a basis model is still least squares, describe how a piecewise polynomial fails and by roughly how much, and count degrees of freedom before and after the constraints.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.