Skip to content
Kudos AI

Least Squares from the Derivative

Minimising the residual sum of squares in closed form, what the slope formula means, and why R-squared cannot compare models of different sizes.

IntermediateModule 130 min · 120 XP
The residual squares drawn literally, shrinking as the line is tilted into place, then the five residuals summed on screen to exactly zero.

Linear regression is usually presented as a formula to apply. It is better understood as the answer to a calculus question: which line makes the squared errors smallest? Once you ask it that way, the coefficients fall out of two derivatives and there is nothing left to memorise.

The objective

Fit y^i=β0+β1xi\hat y_i = \beta_0 + \beta_1 x_i and measure the misfit by the residual sum of squares:

RSS(β0,β1)=∑i=1n(yi−β0−β1xi)2.\mathrm{RSS}(\beta_0, \beta_1) = \sum_{i=1}^{n}\big(y_i - \beta_0 - \beta_1 x_i\big)^2 .

Squaring does two jobs: it stops positive and negative errors cancelling, and it makes RSS a smooth convex function, so setting derivatives to zero finds the global minimum rather than merely a stationary point.

Differentiating

With respect to the intercept:

∂RSS∂β0=−2∑i=1n(yi−β0−β1xi)=0⟹∑i=1nei=0,\frac{\partial \mathrm{RSS}}{\partial \beta_0} = -2\sum_{i=1}^{n}\big(y_i - \beta_0 - \beta_1 x_i\big) = 0 \quad\Longrightarrow\quad \sum_{i=1}^{n} e_i = 0 ,

writing eie_i for the residual. The intercept condition guarantees the residuals sum to exactly zero - equivalently, that the fitted line passes through the point of means (xˉ,yˉ)(\bar x, \bar y). That is not a coincidence or a convention; it is what that one derivative says.

With respect to the slope, and substituting β0=yˉ−β1xˉ\beta_0 = \bar y - \beta_1 \bar x:

β^1=∑i(xi−xˉ)(yi−yˉ)∑i(xi−xˉ)2,β^0=yˉ−β^1xˉ.\hat\beta_1 = \frac{\sum_i (x_i - \bar x)(y_i - \bar y)}{\sum_i (x_i - \bar x)^2}, \qquad \hat\beta_0 = \bar y - \hat\beta_1 \bar x .

Worked example

Five points: x=(1,2,3,4,5)x = (1,2,3,4,5), y=(2.1,3.9,6.2,7.8,10.1)y = (2.1, 3.9, 6.2, 7.8, 10.1).

xˉ=3.0\bar x = 3.0, yˉ=6.02\bar y = 6.02. Then

Sxy=∑(xi−xˉ)(yi−yˉ)=19.90,Sxx=∑(xi−xˉ)2=10.00,S_{xy} = \sum (x_i - \bar x)(y_i - \bar y) = 19.90 , \qquad S_{xx} = \sum (x_i - \bar x)^2 = 10.00 ,

so

β^1=19.9010.00=1.9900,β^0=6.02−1.99×3=0.0500.\hat\beta_1 = \frac{19.90}{10.00} = 1.9900 , \qquad \hat\beta_0 = 6.02 - 1.99 \times 3 = 0.0500 .

The residuals are (0.06, −0.13, 0.18, −0.21, 0.10)(0.06,\, -0.13,\, 0.18,\, -0.21,\, 0.10) - and they sum to exactly 00, as the intercept derivative promised. Check it yourself; that sum is the fastest way to catch an arithmetic slip.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

R-squared, and its trap

R2=1−RSSTSS,TSS=∑i(yi−yˉ)2.R^2 = 1 - \frac{\mathrm{RSS}}{\mathrm{TSS}}, \qquad \mathrm{TSS} = \sum_i (y_i - \bar y)^2 .

Here TSS=39.7080\mathrm{TSS} = 39.7080 and RSS=0.1070\mathrm{RSS} = 0.1070, giving R2=0.9973R^2 = 0.9973 - the line explains 99.73% of the variation in yy.

Why R2R^2 cannot rank models of different size. Adding any predictor, even pure noise, gives least squares strictly more freedom to reduce RSS. RSS can therefore never increase when a predictor is added, so R2R^2 can never decrease. A metric that always rewards more predictors cannot be used to decide whether a predictor is worth having. Comparisons need something that penalises size - adjusted R2R^2, an information criterion, or a held-out estimate from cross-validation.

See the squares. The figure below uses a different five points, (1,2),(2,4),(3,5),(4,4),(5,5)(1,2), (2,4), (3,5), (4,4), (5,5), which scatter far more than the worked example. Each residual is drawn with its square, so RSS is their total area, set against the TSS that the flat mean leaves. Move Intercept and Slope, or press Snap to the least-squares line: it lands on y^=2.2+0.6x\hat y = 2.2 + 0.6x, where RSS=2.4\mathrm{RSS} = 2.4, TSS=6\mathrm{TSS} = 6 and R2=0.6R^2 = 0.6, and the residuals again sum to exactly 00.

Interactive: the squares that least squares minimises

RSS is the total area of the squares.

012345670123456(x̄, ȳ)this lineflat mean4.056
RSS
4.05
TSS
6
R²
0.325
Sum of residuals
-0.5

The squares cover RSS = 4.05, so R² is 0.325. Tilt and lift the line to shrink the total area, and watch the stacked bar fall toward its floor of 2.40, or snap straight to it.

See it move. Least squares is one member of a larger family: slide the parameter and watch a log-likelihood peak exactly where the estimator lands.

Interactive: climbing the likelihood

A fixed sample of 40 observed event counts.

log-likelihoodMLE 1.550.511.522.533.5lambda0123456
ObservedFitted Poisson(λ)
Chosen λ
0.80
Log-likelihood
-72.9
MLE (sample mean)
1.55

Maximum likelihood asks: which λ makes the observed data most probable? Slide λ and the log-likelihood climbs to a single peak, dead on the sample mean (1.55), and the fitted bars snap onto the observed frequencies there. Every GLM on this site is this same climb, just in more dimensions.

Try it live. Fit a generalised linear model with the same normal-equation step iterated to convergence:

A Poisson GLM by hand (IRLS, numpy)

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Before the quiz

Be able to say what setting ∂RSS/∂β0=0\partial \mathrm{RSS}/\partial \beta_0 = 0 guarantees, and why R2R^2 is the wrong tool for model selection. Full derivation: Linear Regression from First Principles.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.