Linear Regression from First Principles
Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.
Prerequisites: What Is Statistical Learning?
Linear regression is old, simple, and still the right first thing to try. It is also the cleanest place to see the pattern every other supervised method follows: write down a measure of how badly the model fits, then choose the parameters that minimise it. Here that minimisation can be done exactly, in closed form, with calculus you already have.
A. The model
With a single predictor, we assume
where is the intercept and the slope. Together they are the model's coefficients. Estimating them from data gives and , and predictions
The -th residual is , the gap between what we saw and what we predicted.
B. What "best fit" means
We need a single number for how badly a candidate line fits. The standard choice is the residual sum of squares:
Squaring does two things: it makes positive and negative misses count the same, and it penalises one large miss more heavily than several small ones. It is also differentiable everywhere, which is what makes the closed-form solution below possible.
Least squares chooses to minimise RSS.
C. Deriving the coefficients
RSS is a smooth function of two variables, so its minimum is where both partial derivatives vanish.
With respect to the intercept:
Dividing by gives , so
Two things follow immediately: the residuals sum to zero, and the fitted line always passes through .
With respect to the slope:
Substituting and rearranging yields
The numerator is the sample covariance of and (up to a constant) and the denominator the sample variance of . So the slope is how much and move together, scaled by how much moves on its own - which is exactly what a slope ought to be.
D. A complete worked fit
Five observations:
| 1 | 1 | 2 |
| 2 | 2 | 4 |
| 3 | 3 | 5 |
| 4 | 4 | 4 |
| 5 | 5 | 5 |
Step 1 - the means.
Step 2 - the deviation products.
| product | |||
|---|---|---|---|
| sum |
Step 3 - the coefficients.
The fitted line is
Step 4 - fitted values and residuals.
| 1 | 2.8 | 2 | 0.64 | |
| 2 | 3.4 | 4 | 0.36 | |
| 3 | 4.0 | 5 | 1.00 | |
| 4 | 4.6 | 4 | 0.36 | |
| 5 | 5.2 | 5 | 0.04 | |
| 0.0 | 2.40 |
The residuals sum to exactly zero, as the derivation promised, and the line passes through - check row 3.
E. How good is the fit?
RSS alone is uninterpretable: it depends on the units and the sample size. Compare it instead against the total sum of squares, the error of the best possible constant prediction :
Then
The predictor explains 60% of the variability in ; the remaining 40% it does not. lies in for a model fitted with an intercept, and is unitless, which is what makes it comparable across problems.
never decreases when you add a predictor, even a column of pure noise, because the extra freedom can only reduce RSS. It therefore cannot be used to choose between models of different sizes - that is what cross-validation is for.
Interactive: the squares that least squares minimises
RSS is the total area of the squares.
- RSS
- 4.05
- TSS
- 6
- R²
- 0.325
- Sum of residuals
- -0.5
The squares cover RSS = 4.05, so R² is 0.325. Tilt and lift the line to shrink the total area, and watch the stacked bar fall toward its floor of 2.40, or snap straight to it.
F. Verifying every number
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Running it prints the table exactly - slope , intercept , residuals
summing to (to machine precision), , ,
- and sklearn: 2.2 0.6 0.6, an independent confirmation.
G. What the model assumes
Least squares always returns an answer. Whether it is meaningful depends on assumptions worth stating plainly:
- Linearity. The relationship really is a straight line. If it curves, the fit is biased no matter how much data you collect - the failure mode from The Bias-Variance Tradeoff.
- Independent errors. Correlated residuals (time series, clustered data) leave coefficients usable but make standard errors far too small.
- Constant error variance. If spread grows with , the fit over-weights the noisy region.
- Influential points. With , a single outlier moves the line substantially - as LOOCV fold 1 showed, dropping moved the prediction there from to , a shift of .
Plotting residuals against fitted values checks the middle two faster than any formal test.
Key takeaways
- Least squares minimises RSS, and setting both partial derivatives to zero gives closed-form coefficients - no iteration required.
- is the covariance of and over the variance of ; forces the line through .
- The derivation guarantees the residuals sum to zero.
- Our worked fit: , , , .
- never decreases when predictors are added, so it cannot compare models of different sizes.
What's next
Squared error is the wrong target when the response is a category rather than a number - a fitted line will happily predict a probability of . Adapting the linear model to classification is Logistic Regression and Classification.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.