Skip to content
Kudos AI

Spline

A piecewise polynomial joined at chosen points called knots, constrained so that the function and its lower derivatives stay continuous there, giving local flexibility without the wild behaviour of a high-degree polynomial.

Also known as: Regression spline, Cubic spline, Natural spline, Smoothing spline

Understanding Spline

A spline is what you get by fitting a separate low-degree polynomial in each region of a predictor and then insisting the pieces meet properly. Left unconstrained, the pieces disagree at the boundaries: fitting four independent cubics to two hundred simulated points with knots at 40, 50 and 60 produces jumps of roughly 6.3, 5.5 and 6.3 units in the fitted value. Requiring continuity, a continuous first derivative and a continuous second derivative removes those jumps, and each constraint costs exactly one degree of freedom, taking a one-knot piecewise cubic from eight parameters to five.

The practical form is a basis function model, so the fit never leaves ordinary least squares. Starting from x, x² and x³, add one truncated power column (x − ξ)³ for x above the knot and zero otherwise, per knot. Adding such a column changes only the third derivative at that knot, leaving the value, slope and curvature continuous, which is why the constraints need not be imposed separately. A cubic spline with K knots is therefore a regression on K + 4 columns, and because a discontinuity in the third derivative is invisible to the eye, cubic is the conventional choice.

Splines are unreliable at the edges of the data, where a cubic has observations on only one side and is free to swing. A natural spline adds the requirement that the function be linear beyond the outermost knots, which is two constraints at each end and takes a K-knot fit down to K degrees of freedom, noticeably narrowing the confidence bands there. Knots are usually placed at uniform quantiles of the observed predictor, so that resolution follows the data, with the number chosen by cross-validation. Compared with a polynomial at the same degrees of freedom, a spline wins because its flexibility is local: a polynomial raises its degree globally and swings hard in the tails.

A smoothing spline abandons knot selection altogether. It minimises the residual sum of squares plus a penalty λ times the integral of the squared second derivative, a direct measure of roughness that is zero for a straight line. The minimiser turns out to be a natural cubic spline with a knot at every unique observation, shrunk by λ. Its nominal parameter count is therefore n, but the meaningful measure is effective degrees of freedom, which fall from n toward 2 as λ grows; at large λ the fit reproduces the least-squares line exactly. Local regression reaches similar ends by fitting a weighted regression at each target point, governed by a span, and generalized additive models carry the whole idea to several predictors by giving each its own curve and adding them.

How to Calculate

f(x) = β₀ + β₁x + β₂x² + β₃x³ + Σₖ βₖ₊₃ (x − ξₖ)³₊

where

ξₖ
the kth knot, a point where the polynomial pieces join
(x − ξ)³₊
the truncated power basis function: (x − ξ)³ when x exceeds the knot, zero otherwise
K + 4
the number of coefficients, and so the degrees of freedom of a cubic spline with K knots
λ ∫ g″(t)² dt
the roughness penalty of a smoothing spline, which replaces knot selection with a tuning parameter

Example of Spline

Take f(x) = 1 + 2x − 0.05x² + 0.001x³ + 0.004(x − 50)³₊ and difference it either side of the knot at 50. The jumps in the value, first derivative and second derivative come out at 9 × 10⁻⁴, 4 × 10⁻⁵ and 4 × 10⁻⁶, which is numerical noise, while the third derivative jumps by 0.024, exactly 6 × 0.004.

Fitting a cubic spline with knots at 40, 50 and 60 to two hundred simulated points is a least squares problem with seven columns, K + 4 = 3 + 4. Differencing the fitted curve at each knot gives jumps of order 10⁻⁴ or smaller, against jumps of more than six units for the unconstrained piecewise fit on the same data.

On the same data, a smoothing spline with a knot at every point has effective degrees of freedom of 186.2 at λ = 10⁻⁶, falling to 2.06 at λ = 10⁶ and to exactly 2.00 at λ = 10¹⁸, where its residual sum of squares matches the ordinary least-squares line’s 12,850.75 to the cent.

Advantages and Disadvantages

Pros

  • Fits by ordinary least squares, so every linear-model tool for inference still applies.
  • Flexibility is local, so the tails stay stable where polynomial fits do not.
  • A single tuning parameter, the number of knots or the roughness penalty, controls the bias-variance trade-off.

Cons

  • Knot number and placement are extra choices, usually resolved by cross-validation rather than by theory.
  • A natural spline buys boundary stability by assuming the relationship straightens out beyond the data.
  • With many predictors the basis grows quickly, and additive models exclude interactions unless they are added explicitly.

Frequently Asked Questions

Why cubic rather than some other degree?

Because a cubic spline is continuous in value, slope and curvature, and only the third derivative jumps at a knot. That discontinuity is not perceptible, so a cubic is the lowest degree that looks completely smooth. Nothing prevents other degrees: a linear spline is continuous with corners allowed, and a piecewise constant fit is a degree-zero spline.

How is a smoothing spline different from a regression spline?

A regression spline takes a small set of knots you choose and fits by least squares. A smoothing spline puts a knot at every unique observation and controls flexibility with a roughness penalty instead. The second removes the knot-placement decision and replaces it with tuning λ, usually by cross-validation.

When should I prefer a GAM to a fully flexible method?

When you need to explain the fit. Because the contributions add, each predictor’s curve can be plotted and read on its own, holding the others fixed. Give that up and methods like boosting or a kernel SVM can capture interactions a GAM cannot, but you lose the ability to say what any single predictor is doing.

The Bottom Line

A spline keeps least squares and changes the columns: piecewise polynomials tied together at knots, with one truncated power column per knot enforcing smoothness for free. It buys local flexibility without a polynomial’s wild tails, and a smoothing spline goes further, replacing knot choice with a roughness penalty whose effective degrees of freedom run from n down to a straight line.