Moving Beyond Linearity: Splines and Additive Models
How to fit curved relationships without leaving least squares: basis functions, the constraints that turn a broken piecewise polynomial into a spline, the single extra column per knot that enforces them for free, and the roughness penalty that lets a curve choose its own flexibility.
Prerequisites: Regularization: Ridge and Lasso
Linear regression assumes a predictor acts in a straight line. When it does not, the instinct is to reach for a more flexible algorithm. There is a cheaper move that keeps least squares and everything built on it, and only changes what you regress on.
A. Basis functions
Pick transformations of that are fixed and known in advance, , and fit
Polynomial regression is the case . Step functions are the case of bin indicators, useful where genuine breakpoints exist and poor otherwise, since a bin cannot show a trend inside itself. The important word is fixed: the transformations are chosen, not estimated, so this is a linear model in the new columns. Standard errors, -statistics and confidence intervals all carry over.
B. Piecewise polynomials break, and constraints fix them
Instead of one global polynomial, fit a cubic in each region, split at knots. With one knot that is two cubics of four parameters each: eight degrees of freedom. Nothing forces the pieces to meet, and they do not. On 200 simulated points with knots at 40, 50 and 60, the fitted value jumps by 6.34, 5.55 and 6.26 units respectively - not rounding, a genuinely discontinuous curve.
The fix is constraints: continuity, then a continuous first derivative, then a continuous second. Each removes exactly one parameter, so , and the result is a cubic spline. In general a cubic spline with knots uses degrees of freedom, since each new knot adds a cubic and three constraints.
C. The column that enforces smoothness for free
Constrained least squares is avoidable. Start from and add one truncated power column per knot,
and the constraints enforce themselves. The claim is checkable. For , differencing across the knot gives
The first three are numerical-differencing error. The last is real and equals exactly. A cubic spline is smooth to the second derivative and breaks only at the third - the one break the eye cannot detect.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Differencing the fitted curve at each knot gives jumps of order or smaller, against more than six units for the unconstrained fit on identical data and knots. Only the columns changed.
Splines still misbehave at the edges, where a cubic has data on one side only. A natural spline requires linearity beyond the outermost knots - two constraints at each end - taking degrees of freedom down to and visibly narrowing the boundary confidence bands. And this is why splines beat polynomials at equal flexibility: a polynomial raises its degree globally and swings in the tails, while a spline adds a basis function that is zero until its knot.
The figure below fits both bases to one sample, and the basis is a switch. With the truncated power columns in place the curve is continuous in f, f′ and f′′ at every knot - and those are exact zeros, not small numbers, because the basis cannot express a break there for least squares to find. Drop to a free cubic per region and the same data, the same knots and the same least squares tear the curve open. Slide the knots anywhere you like; the smoothness does not depend on where they are.
Interactive: same data, same knots, different columns
Switch the basis and watch the curve tear at the knots.
- Columns
- 7
- Residual sum of squares
- 11844
- Worst jump in f
- 0
- Worst jump in f′′′
- 0.129
7 columns for 3 knots - the K + 4 of the degree-of-freedom count, now visible as a design matrix. The curve is continuous in f, f′ and f′′ at every knot, and those are exact zeros rather than small numbers: the basis cannot express a discontinuity there, so least squares could not fit one if the data begged for it. All that is left to break is the third derivative, which is the break no one can see. Take the truncated power columns away and look at the same fit.
D. Penalise roughness instead of choosing knots
A smoothing spline removes knot selection entirely. Minimise
The second derivative measures roughness - zero for a straight line - so the integral totals the wiggle. The minimiser is a natural cubic spline with a knot at every unique observation, shrunk by .
That sounds like parameters and certain overfitting, and the resolution is effective degrees of freedom, the trace of the matrix mapping to fitted values. Measured on the same 200 points:
At the bottom the fit has 2.00 effective degrees of freedom and a residual sum of squares of , the same to the cent as the ordinary least-squares line. The limit is not approximately a straight line; it is the straight line.
The other end needs one honest remark. At exactly the penalty vanishes, and any curve through all 200 points has zero RSS, so the objective alone does not pick one. The natural spline does: with a knot at each of the 200 distinct it has exactly 200 columns for 200 rows, and the fit is the unique natural cubic interpolant. It is also the limit as , where effective degrees of freedom rise to exactly .
Local regression reaches similar ends differently: at each target point, take the nearest fraction of the data, weight by distance, and fit a weighted line. The span plays 's role, and the method is memory-based, since every prediction refits.
E. Many predictors, one curve each
A generalized additive model gives each predictor its own function and adds them:
Because they add, each can be plotted and read as that predictor's effect holding the others fixed. That interpretability is the reason to prefer a GAM to boosting or a kernel SVM. The cost is in the same word: an effect that exists only in a combination of predictors is invisible unless you add the interaction by hand.
With natural splines as components the model is one large least squares fit. With smoothing splines it is fitted by backfitting, cycling through the predictors and re-fitting each against the partial residuals until nothing changes.
Where this leaves you
You do not need a new algorithm to fit a curve. Transform the columns and least squares still applies. Constraints buy smoothness at exactly one degree of freedom each, and a well-chosen basis makes them free. A penalty replaces knot selection with a single tuning parameter whose limits are an interpolant and a straight line. And additivity carries all of it into many predictors while keeping the one property flexible methods give up: you can still say what each predictor is doing. The training path Moving Beyond Linearity works each of these by hand and in code.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.