Logistic Regression and Classification
Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.
Prerequisites: Linear Regression from First Principles, Bayes' Theorem and Belief Updating
Linear Regression from First Principles fitted a numeric response. Many problems instead ask a categorical question: does this belong to class A or class B? What we usually want is not the label but the probability of the label, and probabilities are constrained in a way that straight lines are not.
A. Why not just fit a line?
Code the two classes as and and run least squares. The fitted line is , defined for every real , and therefore eventually exceeds and eventually falls below . As an estimate of that is not merely inaccurate, it is impossible - it violates the axioms from Probability from Zero.
We need a model whose output is confined to by construction.
B. The logistic function
Logistic regression models the probability directly:
Writing , this is the logistic (or sigmoid) function
the two forms being algebraically identical (divide numerator and denominator of the first by ). Its behaviour is exactly what we need: as , ; as , ; and . The output can approach the endpoints but never reach them.
A few values, with :
| 2 | 0.1192 | |
| 4 | 0.5000 | |
| 6 | 0.8808 |
C. Odds and log-odds: what the coefficients mean
Rearranging the model gives the odds:
and taking logs gives the log-odds or logit:
This is the sense in which logistic regression is a linear model - linear in the log-odds, not in the probability. Odds run from to and log-odds over the whole real line, which is precisely the range a linear function needs.
| odds | log-odds | |
|---|---|---|
| 0.20 | 0.25 | |
| 0.50 | 1.00 | |
| 0.75 | 3.00 | |
| 0.90 | 9.00 |
So is the change in log-odds per one-unit increase in , and is the multiplicative change in the odds.
The most common misreading. is not the change in probability per unit of . Because the logistic curve is S-shaped, the same one-unit step moves the probability a lot near and almost not at all out in the tails. Only the log-odds change by a constant amount.
D. Fitting by maximum likelihood
Least squares is not the natural criterion here. Instead we choose the coefficients that make the observed labels most probable. Each observation contributes if and if , which combines into the likelihood
Taking logs turns the product into a sum - numerically far better behaved - giving the log-likelihood
Unlike least squares, this has no closed-form solution; it is maximised numerically. Its gradient is remarkably clean:
Each observation pushes the coefficients in proportion to how wrong the current prediction is - a pattern that reappears in Backpropagation and Gradient Descent.
E. One gradient step, by hand
Six observations, deliberately not perfectly separable (the reason is in section G):
| 1 | 2 | 3 | 4 | 5 | 6 | |
|---|---|---|---|---|---|---|
| 0 | 1 | 0 | 1 | 1 | 1 |
Start from . Then and for every observation, so
The gradient components:
Taking a step of size uphill (we are maximising):
The new probabilities are and the log-likelihood has risen to . One step improved the fit.
Step size is not a free parameter. With the same step lands at and the log-likelihood falls to - far worse than where it started. Overshooting a maximum is easy; that a step is uphill locally does not mean any step size along it is an improvement.
Iterating to convergence gives
with log-likelihood .
Interpreting the result. Each extra unit of multiplies the odds by - roughly tripling them. The decision boundary, where , is where :
Below we predict class 0, above it class 1.
Interactive: linear in the log-odds, curved in the probability
The marked step is one unit of x, drawn on both.
- Odds multiplier
- 3.141x
- Boundary (p = 0.5)
- 2.4199
- Step adds, in log-odds
- 1.1447
- Step adds, in probability
- 0.2642
The step adds 1.1447 to the log-odds - the same amount wherever you take it, because that line is straight - and multiplies the odds by 3.141, also everywhere. But it moves the probability by 0.2642, and that is near its largest, because you are close to the boundary where the curve is steepest. Slide out into a tail and watch the same step become worth almost nothing.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Every number above is reproduced by running this.
F. From probability to decision
The model outputs a probability; a decision needs a threshold. The default is only correct when the two kinds of mistake cost the same. When a false negative is far more costly than a false positive, the right threshold is lower - and this is a decision about consequences, not a statistical question the data can answer.
The base-rate lesson from Bayes' Theorem applies here in full: with a rare positive class, a model can be highly accurate and still be wrong most of the times it predicts "positive".
G. When the fit blows up
If the classes are perfectly separable - some threshold splits them with no overlap - the likelihood can always be increased by making larger, pushing fitted probabilities towards exactly and . There is no maximum, and the MLE does not exist.
This is not hypothetical. Replacing our data with the cleanly separated , makes an unregularised solver return with - numbers that are artefacts of where the optimiser stopped, not estimates. Software may or may not warn you. The standard remedy is a penalty on coefficient size, which is exactly the subject of the next article.
Key takeaways
- A linear model of a probability produces values outside ; the logistic function confines the output to by construction.
- Logistic regression is linear in the log-odds: is the change in log-odds per unit of , and the odds multiplier - not a change in probability.
- Fitting maximises the log-likelihood, which has no closed form and is solved numerically; the gradient is and .
- Our fit converged to , : odds per unit, boundary at .
- Step size matters - improved the fit, made it much worse.
- Under perfect separation the MLE does not exist and coefficients diverge.
What's next
Both failure modes just seen - coefficients running away under separation, and flexible models chasing noise - are treated by the same idea: add a penalty that makes large coefficients expensive. That is Regularization: Ridge and Lasso.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.