Skip to content
Kudos AI
Lire en français
Supervised Learning

Logistic Regression and Classification

Why a straight line cannot model a probability, how the logistic function fixes it, and what the coefficients mean in log-odds, with a gradient-ascent step and a converged fit computed and checked numerically.

7 min readKudos AI

Prerequisites: Linear Regression from First Principles, Bayes' Theorem and Belief Updating

A straight line running out past 1 and below 0, the same linear part folded into the logistic curve, and the boundary read straight off the coefficients.

Linear Regression from First Principles fitted a numeric response. Many problems instead ask a categorical question: does this belong to class A or class B? What we usually want is not the label but the probability of the label, and probabilities are constrained in a way that straight lines are not.

A. Why not just fit a line?

Code the two classes as 00 and 11 and run least squares. The fitted line is β^0+β^1X\hat\beta_0 + \hat\beta_1 X, defined for every real XX, and therefore eventually exceeds 11 and eventually falls below 00. As an estimate of P(Y=1∣X)P(Y = 1 \mid X) that is not merely inaccurate, it is impossible - it violates the axioms from Probability from Zero.

We need a model whose output is confined to (0,1)(0, 1) by construction.

B. The logistic function

Logistic regression models the probability directly:

p(X)=P(Y=1∣X)=eβ0+β1X1+eβ0+β1X.p(X) = P(Y = 1 \mid X) = \frac{e^{\beta_0 + \beta_1 X}}{1 + e^{\beta_0 + \beta_1 X}} .

Writing z=β0+β1Xz = \beta_0 + \beta_1 X, this is the logistic (or sigmoid) function

σ(z)=11+e−z,\sigma(z) = \frac{1}{1 + e^{-z}} ,

the two forms being algebraically identical (divide numerator and denominator of the first by eze^{z}). Its behaviour is exactly what we need: as z→+∞z \to +\infty, σ→1\sigma \to 1; as z→−∞z \to -\infty, σ→0\sigma \to 0; and σ(0)=0.5\sigma(0) = 0.5. The output can approach the endpoints but never reach them.

A few values, with z=−4+xz = -4 + x:

xxzzσ(z)\sigma(z)
2−2-20.1192
4000.5000
6220.8808

C. Odds and log-odds: what the coefficients mean

Rearranging the model gives the odds:

p(X)1−p(X)=eβ0+β1X,\frac{p(X)}{1 - p(X)} = e^{\beta_0 + \beta_1 X} ,

and taking logs gives the log-odds or logit:

log⁡ ⁣(p(X)1−p(X))=β0+β1X.\log\!\left(\frac{p(X)}{1 - p(X)}\right) = \beta_0 + \beta_1 X .

This is the sense in which logistic regression is a linear model - linear in the log-odds, not in the probability. Odds run from 00 to ∞\infty and log-odds over the whole real line, which is precisely the range a linear function needs.

ppodds p/(1−p)p/(1-p)log-odds
0.200.25−1.3863-1.3863
0.501.000.00000.0000
0.753.001.09861.0986
0.909.002.19722.1972

So β1\beta_1 is the change in log-odds per one-unit increase in XX, and eβ1e^{\beta_1} is the multiplicative change in the odds.

The most common misreading. β1\beta_1 is not the change in probability per unit of XX. Because the logistic curve is S-shaped, the same one-unit step moves the probability a lot near p=0.5p = 0.5 and almost not at all out in the tails. Only the log-odds change by a constant amount.

D. Fitting by maximum likelihood

Least squares is not the natural criterion here. Instead we choose the coefficients that make the observed labels most probable. Each observation contributes p(xi)p(x_i) if yi=1y_i = 1 and 1−p(xi)1 - p(x_i) if yi=0y_i = 0, which combines into the likelihood

ℓ(β0,β1)=∏i: yi=1p(xi)∏i: yi=0(1−p(xi)).\ell(\beta_0, \beta_1) = \prod_{i:\,y_i=1} p(x_i) \prod_{i:\,y_i=0}\big(1 - p(x_i)\big) .

Taking logs turns the product into a sum - numerically far better behaved - giving the log-likelihood

log⁡ℓ=∑i=1n[yilog⁡p(xi)+(1−yi)log⁡(1−p(xi))].\log \ell = \sum_{i=1}^{n}\Big[y_i \log p(x_i) + (1-y_i)\log\big(1 - p(x_i)\big)\Big] .

Unlike least squares, this has no closed-form solution; it is maximised numerically. Its gradient is remarkably clean:

∂log⁡ℓ∂β0=∑i(yi−pi),∂log⁡ℓ∂β1=∑i(yi−pi) xi.\frac{\partial \log\ell}{\partial \beta_0} = \sum_i (y_i - p_i), \qquad \frac{\partial \log\ell}{\partial \beta_1} = \sum_i (y_i - p_i)\,x_i .

Each observation pushes the coefficients in proportion to how wrong the current prediction is - a pattern that reappears in Backpropagation and Gradient Descent.

E. One gradient step, by hand

Six observations, deliberately not perfectly separable (the reason is in section G):

xix_i123456
yiy_i010111

Start from β0=β1=0\beta_0 = \beta_1 = 0. Then zi=0z_i = 0 and pi=σ(0)=0.5p_i = \sigma(0) = 0.5 for every observation, so

log⁡ℓ=6log⁡(0.5)=−4.158883.\log\ell = 6\log(0.5) = -4.158883 .

The gradient components:

∑i(yi−pi)=(0−0.5)+(1−0.5)+(0−0.5)+(1−0.5)+(1−0.5)+(1−0.5)=1.0,\sum_i (y_i - p_i) = (0-0.5)+(1-0.5)+(0-0.5)+(1-0.5)+(1-0.5)+(1-0.5) = 1.0 , ∑i(yi−pi)xi=−0.5+1.0−1.5+2.0+2.5+3.0=6.5.\sum_i (y_i - p_i)x_i = -0.5 + 1.0 - 1.5 + 2.0 + 2.5 + 3.0 = 6.5 .

Taking a step of size η=0.05\eta = 0.05 uphill (we are maximising):

β0←0+0.05(1.0)=0.05,β1←0+0.05(6.5)=0.325.\beta_0 \leftarrow 0 + 0.05(1.0) = 0.05, \qquad \beta_1 \leftarrow 0 + 0.05(6.5) = 0.325 .

The new probabilities are [0.5927,0.6682,0.7359,0.7941,0.8422,0.8808][0.5927, 0.6682, 0.7359, 0.7941, 0.8422, 0.8808] and the log-likelihood has risen to −3.162034-3.162034. One step improved the fit.

Step size is not a free parameter. With η=0.5\eta = 0.5 the same step lands at β=(0.5,3.25)\beta = (0.5, 3.25) and the log-likelihood falls to −14.024-14.024 - far worse than where it started. Overshooting a maximum is easy; that a step is uphill locally does not mean any step size along it is an improvement.

Iterating to convergence gives

β^0=−2.7700,β^1=1.1447,\hat\beta_0 = -2.7700, \qquad \hat\beta_1 = 1.1447,

with log-likelihood −2.440125-2.440125.

Interpreting the result. Each extra unit of xx multiplies the odds by e1.14466=3.1414e^{1.14466} = 3.1414 - roughly tripling them. The decision boundary, where p=0.5p = 0.5, is where z=0z = 0:

x=−β^0β^1=2.77001.14466=2.4199.x = -\frac{\hat\beta_0}{\hat\beta_1} = \frac{2.7700}{1.14466} = 2.4199 .

Below x≈2.42x \approx 2.42 we predict class 0, above it class 1.

Interactive: linear in the log-odds, curved in the probability

The marked step is one unit of x, drawn on both.

00.5
Odds multiplier
3.141x
Boundary (p = 0.5)
2.4199
Step adds, in log-odds
1.1447
Step adds, in probability
0.2642

The step adds 1.1447 to the log-odds - the same amount wherever you take it, because that line is straight - and multiplies the odds by 3.141, also everywhere. But it moves the probability by 0.2642, and that is near its largest, because you are close to the boundary where the curve is steepest. Slide out into a tail and watch the same step become worth almost nothing.

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Every number above is reproduced by running this.

F. From probability to decision

The model outputs a probability; a decision needs a threshold. The default 0.50.5 is only correct when the two kinds of mistake cost the same. When a false negative is far more costly than a false positive, the right threshold is lower - and this is a decision about consequences, not a statistical question the data can answer.

The base-rate lesson from Bayes' Theorem applies here in full: with a rare positive class, a model can be highly accurate and still be wrong most of the times it predicts "positive".

G. When the fit blows up

If the classes are perfectly separable - some threshold splits them with no overlap - the likelihood can always be increased by making β1\beta_1 larger, pushing fitted probabilities towards exactly 00 and 11. There is no maximum, and the MLE does not exist.

This is not hypothetical. Replacing our data with the cleanly separated x=(1,2,3,4)x = (1,2,3,4), y=(0,0,1,1)y = (0,0,1,1) makes an unregularised solver return β^1≈18.2\hat\beta_1 \approx 18.2 with β^0≈−45.8\hat\beta_0 \approx -45.8 - numbers that are artefacts of where the optimiser stopped, not estimates. Software may or may not warn you. The standard remedy is a penalty on coefficient size, which is exactly the subject of the next article.

Key takeaways

  • A linear model of a probability produces values outside [0,1][0,1]; the logistic function confines the output to (0,1)(0,1) by construction.
  • Logistic regression is linear in the log-odds: β1\beta_1 is the change in log-odds per unit of XX, and eβ1e^{\beta_1} the odds multiplier - not a change in probability.
  • Fitting maximises the log-likelihood, which has no closed form and is solved numerically; the gradient is ∑(yi−pi)\sum(y_i - p_i) and ∑(yi−pi)xi\sum(y_i - p_i)x_i.
  • Our fit converged to β^0=−2.7700\hat\beta_0 = -2.7700, β^1=1.1447\hat\beta_1 = 1.1447: odds ×3.14\times 3.14 per unit, boundary at x=2.4199x = 2.4199.
  • Step size matters - η=0.05\eta = 0.05 improved the fit, η=0.5\eta = 0.5 made it much worse.
  • Under perfect separation the MLE does not exist and coefficients diverge.

What's next

Both failure modes just seen - coefficients running away under separation, and flexible models chasing noise - are treated by the same idea: add a penalty that makes large coefficients expensive. That is Regularization: Ridge and Lasso.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

10 min readSupervised Learning

Comparing Classifiers, and What Accuracy Hides

The Bayes classifier nothing can beat and the error floor it leaves behind, k-nearest-neighbours as a nonparametric imitation with k as the flexibility dial, discriminant analysis and why a shared covariance forces a straight line, and the confusion matrix, thresholds and ROC curve that a single accuracy figure conceals - every number computed on simulated data where the optimum is known.

Machine LearningStatistics
7 min readSupport Vector Machines

Support Vector Machines: Margins and Kernels

Why the widest slab between two classes is a good boundary, why insisting on a perfect one is self-defeating, how a budget for violations buys back stability, and how a kernel bends the boundary by working in a space it never has to build.

Machine LearningOptimization
7 min readSupervised Learning

Linear Regression from First Principles

Derive the least-squares coefficients by differentiating the residual sum of squares, then work a complete five-observation fit by hand: coefficients, fitted values, residuals, RSS, and R-squared, each verified numerically.

StatisticsMachine LearningMathematics
← Back to all articles