Least Squares from the Derivative
Minimising the residual sum of squares in closed form, what the slope formula means, and why R-squared cannot compare models of different sizes.
Linear regression is usually presented as a formula to apply. It is better understood as the answer to a calculus question: which line makes the squared errors smallest? Once you ask it that way, the coefficients fall out of two derivatives and there is nothing left to memorise.
The objective
Fit and measure the misfit by the residual sum of squares:
Squaring does two jobs: it stops positive and negative errors cancelling, and it makes RSS a smooth convex function, so setting derivatives to zero finds the global minimum rather than merely a stationary point.
Differentiating
With respect to the intercept:
writing for the residual. The intercept condition guarantees the residuals sum to exactly zero - equivalently, that the fitted line passes through the point of means . That is not a coincidence or a convention; it is what that one derivative says.
With respect to the slope, and substituting :
Worked example
Five points: , .
, . Then
so
The residuals are - and they sum to exactly , as the intercept derivative promised. Check it yourself; that sum is the fastest way to catch an arithmetic slip.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
R-squared, and its trap
Here and , giving - the line explains 99.73% of the variation in .
Why cannot rank models of different size. Adding any predictor, even pure noise, gives least squares strictly more freedom to reduce RSS. RSS can therefore never increase when a predictor is added, so can never decrease. A metric that always rewards more predictors cannot be used to decide whether a predictor is worth having. Comparisons need something that penalises size - adjusted , an information criterion, or a held-out estimate from cross-validation.
See the squares. The figure below uses a different five points, , which scatter far more than the worked example. Each residual is drawn with its square, so RSS is their total area, set against the TSS that the flat mean leaves. Move Intercept and Slope, or press Snap to the least-squares line: it lands on , where , and , and the residuals again sum to exactly .
Interactive: the squares that least squares minimises
RSS is the total area of the squares.
- RSS
- 4.05
- TSS
- 6
- R²
- 0.325
- Sum of residuals
- -0.5
The squares cover RSS = 4.05, so R² is 0.325. Tilt and lift the line to shrink the total area, and watch the stacked bar fall toward its floor of 2.40, or snap straight to it.
See it move. Least squares is one member of a larger family: slide the parameter and watch a log-likelihood peak exactly where the estimator lands.
Interactive: climbing the likelihood
A fixed sample of 40 observed event counts.
- Chosen λ
- 0.80
- Log-likelihood
- -72.9
- MLE (sample mean)
- 1.55
Maximum likelihood asks: which λ makes the observed data most probable? Slide λ and the log-likelihood climbs to a single peak, dead on the sample mean (1.55), and the fitted bars snap onto the observed frequencies there. Every GLM on this site is this same climb, just in more dimensions.
Try it live. Fit a generalised linear model with the same normal-equation step iterated to convergence:
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Before the quiz
Be able to say what setting guarantees, and why is the wrong tool for model selection. Full derivation: Linear Regression from First Principles.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.