Skip to content
Kudos AI

Linear Regression

A model that predicts a numeric response as a weighted sum of the predictors, fitted by minimizing squared error.

Also known as: Ordinary least squares, OLS

The residual squares drawn literally, shrinking as the line is tilted into place, then the five residuals summed on screen to exactly zero.

Understanding Linear Regression

Linear regression models the response as an intercept plus a weighted sum of predictors. Fitting means choosing the weights that minimize the sum of squared residuals, the vertical gaps between observed and predicted values. Squaring makes positive and negative errors both count and penalizes large misses disproportionately.

The estimate has a closed form obtained by setting the derivatives of the squared-error criterion to zero, which yields the normal equations. This is a genuine practical advantage: the fit is exact and immediate rather than approached iteratively, and there is no learning rate or convergence behaviour to worry about.

The interpretation of a coefficient is precise and easily misstated. It is the expected change in the response for a one-unit increase in that predictor with all other predictors held constant. That "holding constant" clause is doing real work. When predictors are correlated, the data contains little information about what happens when one moves and the others do not, and individual coefficients become unstable and hard to interpret even while the model as a whole predicts well.

The word "linear" constrains the coefficients, not the predictors. Adding a squared term, a logarithm, or a product of two predictors as a new column keeps the model linear in its parameters, so curved and interacting relationships are entirely within reach. What linear regression genuinely cannot do is learn such a transformation on its own.

How to Calculate

ŷ = β₀ + β₁x₁ + … + βₚxₚ, fitted by minimizing Σᵢ (yᵢ − ŷᵢ)²

where

ŷ
the predicted response
β₀
the intercept, the prediction when every predictor is zero
βⱼ
the coefficient on predictor j
yᵢ − ŷᵢ
the residual for observation i

Example of Linear Regression

Predicting house price from floor area might give price = 50,000 + 1,200 × area. The intercept is the fitted value at zero area, which has no physical meaning here and should not be over-interpreted; the slope says each additional square metre is associated with about 1,200 more in price.

Adding a second predictor, say distance from the city centre, changes the meaning of the first coefficient. It now estimates the effect of area among houses at the same distance. If larger houses tend to sit further out, the two coefficients will differ substantially from what either predictor gives on its own.

The association is not causal. The coefficient describes how price and area vary together in this sample, given the other predictors in the model. Whether enlarging a house would raise its price by 1,200 per square metre is a different question that the regression alone cannot answer.

Advantages and Disadvantages

Pros

  • Highly interpretable: every coefficient has a direct, stateable meaning.
  • Fits in closed form, with no hyperparameters to tune.
  • Comes with well-developed inference: standard errors, confidence intervals, and tests.

Cons

  • Cannot represent non-linear structure unless it is supplied manually as new features.
  • Sensitive to outliers, because squaring magnifies large residuals.
  • Correlated predictors make individual coefficients unstable and difficult to interpret.

Frequently Asked Questions

What does R² actually measure?

The proportion of variance in the response explained by the model. It never decreases when predictors are added, even useless ones, so it cannot be used to compare models of different sizes; adjusted R² or a cross-validated error estimate should be used instead.

Does linear regression require normally distributed data?

The predictors and the response need not be normal. Normality of the residuals is an assumption of the classical inference procedures, the t-tests and confidence intervals, not of the least-squares fit itself, which remains well defined regardless.

What should be done about correlated predictors?

Options include dropping redundant predictors, combining them, or applying ridge regression, which stabilizes the coefficients by penalizing their magnitude. Prediction accuracy often survives collinearity intact; it is coefficient interpretation that suffers.

The Bottom Line

Linear regression predicts a numeric outcome as a weighted sum, fits in closed form, and yields coefficients that can be stated in plain language. Its limits are non-linearity, outliers, and correlated predictors, and it remains the baseline that more complex models should be required to beat.