Skip to content
Kudos AI
Lire en français
Supervised Learning

Moving Beyond Linearity: Splines and Additive Models

How to fit curved relationships without leaving least squares: basis functions, the constraints that turn a broken piecewise polynomial into a spline, the single extra column per knot that enforces them for free, and the roughness penalty that lets a curve choose its own flexibility.

6 min readKudos AI

Prerequisites: Regularization: Ridge and Lasso

A straight line failing on curved data, then the same least squares fit run on transformed columns, and finally two cubics meeting at a knot with a visible jump that constraints close.

Linear regression assumes a predictor acts in a straight line. When it does not, the instinct is to reach for a more flexible algorithm. There is a cheaper move that keeps least squares and everything built on it, and only changes what you regress on.

A. Basis functions

Pick transformations of XX that are fixed and known in advance, b1(X),…,bK(X)b_1(X), \dots, b_K(X), and fit

yi=β0+β1b1(xi)+⋯+βKbK(xi)+εi.y_i = \beta_0 + \beta_1 b_1(x_i) + \cdots + \beta_K b_K(x_i) + \varepsilon_i .

Polynomial regression is the case bj(x)=xjb_j(x) = x^j. Step functions are the case of bin indicators, useful where genuine breakpoints exist and poor otherwise, since a bin cannot show a trend inside itself. The important word is fixed: the transformations are chosen, not estimated, so this is a linear model in the new columns. Standard errors, FF-statistics and confidence intervals all carry over.

B. Piecewise polynomials break, and constraints fix them

Instead of one global polynomial, fit a cubic in each region, split at knots. With one knot that is two cubics of four parameters each: eight degrees of freedom. Nothing forces the pieces to meet, and they do not. On 200 simulated points with knots at 40, 50 and 60, the fitted value jumps by 6.34, 5.55 and 6.26 units respectively - not rounding, a genuinely discontinuous curve.

The fix is constraints: continuity, then a continuous first derivative, then a continuous second. Each removes exactly one parameter, so 8−3=58 - 3 = 5, and the result is a cubic spline. In general a cubic spline with KK knots uses K+4K + 4 degrees of freedom, since each new knot adds a cubic and three constraints.

C. The column that enforces smoothness for free

Constrained least squares is avoidable. Start from x,x2,x3x, x^2, x^3 and add one truncated power column per knot,

h(x,ξ)=(x−ξ)+3={(x−ξ)3x>ξ0otherwise,h(x, \xi) = (x - \xi)_+^3 = \begin{cases} (x - \xi)^3 & x > \xi \\ 0 & \text{otherwise,} \end{cases}

and the constraints enforce themselves. The claim is checkable. For f(x)=1+2x−0.05x2+0.001x3+0.004(x−50)+3f(x) = 1 + 2x - 0.05x^2 + 0.001x^3 + 0.004(x-50)_+^3, differencing across the knot gives

leftrightjumpf100.9996101.00059×10−4f′4.499984.500024×10−5f′′0.2000060.2000104×10−6f′′′0.0060.0300.024\begin{array}{lrrl} & \text{left} & \text{right} & \text{jump} \\ \hline f & 100.9996 & 101.0005 & 9 \times 10^{-4} \\ f' & 4.49998 & 4.50002 & 4 \times 10^{-5} \\ f'' & 0.200006 & 0.200010 & 4 \times 10^{-6} \\ f''' & 0.006 & 0.030 & 0.024 \end{array}

The first three are numerical-differencing error. The last is real and equals 6β46\beta_4 exactly. A cubic spline is smooth to the second derivative and breaks only at the third - the one break the eye cannot detect.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Differencing the fitted curve at each knot gives jumps of order 10−410^{-4} or smaller, against more than six units for the unconstrained fit on identical data and knots. Only the columns changed.

Splines still misbehave at the edges, where a cubic has data on one side only. A natural spline requires linearity beyond the outermost knots - two constraints at each end - taking K+4K + 4 degrees of freedom down to KK and visibly narrowing the boundary confidence bands. And this is why splines beat polynomials at equal flexibility: a polynomial raises its degree globally and swings in the tails, while a spline adds a basis function that is zero until its knot.

The figure below fits both bases to one sample, and the basis is a switch. With the truncated power columns in place the curve is continuous in f, f′ and f′′ at every knot - and those are exact zeros, not small numbers, because the basis cannot express a break there for least squares to find. Drop to a free cubic per region and the same data, the same knots and the same least squares tear the curve open. Slide the knots anywhere you like; the smoothness does not depend on where they are.

Interactive: same data, same knots, different columns

Switch the basis and watch the curve tear at the knots.

334558
Columns
7
Residual sum of squares
11844
Worst jump in f
0
Worst jump in f′′′
0.129

7 columns for 3 knots - the K + 4 of the degree-of-freedom count, now visible as a design matrix. The curve is continuous in f, f′ and f′′ at every knot, and those are exact zeros rather than small numbers: the basis cannot express a discontinuity there, so least squares could not fit one if the data begged for it. All that is left to break is the third derivative, which is the break no one can see. Take the truncated power columns away and look at the same fit.

D. Penalise roughness instead of choosing knots

A smoothing spline removes knot selection entirely. Minimise

∑i=1n(yi−g(xi))2+λ∫g′′(t)2 dt.\sum_{i=1}^{n}\big(y_i - g(x_i)\big)^2 + \lambda \int g''(t)^2\,\mathrm{d}t .

The second derivative measures roughness - zero for a straight line - so the integral totals the wiggle. The minimiser is a natural cubic spline with a knot at every unique observation, shrunk by λ\lambda.

That sounds like nn parameters and certain overfitting, and the resolution is effective degrees of freedom, the trace of the matrix mapping yy to fitted values. Measured on the same 200 points:

λRSSeffective df10−61,021186.210010,21325.210612,8012.06101812,850.752.00\begin{array}{rrr} \lambda & \text{RSS} & \text{effective df} \\ \hline 10^{-6} & 1{,}021 & 186.2 \\ 10^{0} & 10{,}213 & 25.2 \\ 10^{6} & 12{,}801 & 2.06 \\ 10^{18} & 12{,}850.75 & 2.00 \end{array}

At the bottom the fit has 2.00 effective degrees of freedom and a residual sum of squares of 12,850.7512{,}850.75, the same to the cent as the ordinary least-squares line. The limit is not approximately a straight line; it is the straight line.

The other end needs one honest remark. At λ=0\lambda = 0 exactly the penalty vanishes, and any curve through all 200 points has zero RSS, so the objective alone does not pick one. The natural spline does: with a knot at each of the 200 distinct xx it has exactly 200 columns for 200 rows, and the fit is the unique natural cubic interpolant. It is also the limit as λ→0\lambda \to 0, where effective degrees of freedom rise to exactly n=200n = 200.

Local regression reaches similar ends differently: at each target point, take the nearest fraction ss of the data, weight by distance, and fit a weighted line. The span ss plays λ\lambda's role, and the method is memory-based, since every prediction refits.

E. Many predictors, one curve each

A generalized additive model gives each predictor its own function and adds them:

yi=β0+f1(xi1)+⋯+fp(xip)+εi.y_i = \beta_0 + f_1(x_{i1}) + \cdots + f_p(x_{ip}) + \varepsilon_i .

Because they add, each fjf_j can be plotted and read as that predictor's effect holding the others fixed. That interpretability is the reason to prefer a GAM to boosting or a kernel SVM. The cost is in the same word: an effect that exists only in a combination of predictors is invisible unless you add the interaction by hand.

With natural splines as components the model is one large least squares fit. With smoothing splines it is fitted by backfitting, cycling through the predictors and re-fitting each against the partial residuals until nothing changes.

Where this leaves you

You do not need a new algorithm to fit a curve. Transform the columns and least squares still applies. Constraints buy smoothness at exactly one degree of freedom each, and a well-chosen basis makes them free. A penalty replaces knot selection with a single tuning parameter whose limits are an interpolant and a straight line. And additivity carries all of it into many predictors while keeping the one property flexible methods give up: you can still say what each predictor is doing. The training path Moving Beyond Linearity works each of these by hand and in code.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
4 min readTime Series

A Score That Loses to Doing Nothing

A five-nearest-neighbour model scores 0.9983 under random five-fold cross-validation on a random walk, a series whose increments are by construction unpredictable. Evaluated forward in time it scores 0.6559 with an RMSE 12.44 times larger, and loses to carrying the last observed value forward. The split, not the model, produced the first number.

StatisticsMachine Learning
7 min readAnomaly Detection

The Detector That Never Fires Is 99.5% Accurate

At a realistic base rate the do-nothing detector wins on accuracy, a ROC of 0.9468 hides an alert queue that is 64% false, distance from the mean scores below chance when anomalies sit at the centre, and twenty anomalies that group together hide each other from the method built to find them.

Machine LearningStatistics
← Back to all articles