Skip to content
Kudos AI

Logistic Regression

A classification model that predicts the probability of a class by passing a linear combination of predictors through the logistic function.

A straight line running out past 1 and below 0, the same linear part folded into the logistic curve, and the boundary read straight off the coefficients.

Understanding Logistic Regression

Applying linear regression directly to a binary outcome fails for a structural reason: a linear function is unbounded, so it will predict values below 0 and above 1, which cannot be probabilities. Logistic regression keeps the linear combination but passes it through the logistic function, which maps the whole real line into the open interval (0, 1). James and colleagues introduce it in exactly these terms, as the way to model a probability with a function that stays within range for every value of the predictors.

The transformation has an interpretable inverse. Solving for the linear part shows that it equals the log-odds, the logarithm of p/(1 − p). So while the effect of a predictor on the probability is non-linear, its effect on the log-odds is exactly linear, and each coefficient can be read as the change in log-odds per unit increase of its predictor. Exponentiating a coefficient converts it into an odds ratio, which is usually easier to communicate.

Fitting is by maximum likelihood rather than least squares. The likelihood multiplies the predicted probability of the correct outcome across observations, and the coefficients chosen are those maximizing it. There is no closed-form solution, so the optimization is iterative, but the objective is convex, which means there is a single optimum and no risk of converging to a poor local one.

The decision boundary is linear. Classifying at a threshold of 0.5 corresponds to the linear part being zero, which defines a hyperplane in predictor space. Logistic regression therefore cannot separate classes divided by a genuinely curved boundary unless non-linear terms are supplied as features, exactly as with linear regression.

How to Calculate

p(X) = e^(β₀+β₁X) / (1 + e^(β₀+β₁X)), equivalently log(p/(1−p)) = β₀ + β₁X

where

p(X)
the predicted probability of the positive class
β₀, β₁
the intercept and coefficient, estimated by maximum likelihood
p/(1−p)
the odds of the positive class
log(p/(1−p))
the log-odds, or logit, which is linear in X

Example of Logistic Regression

Suppose a model of default risk against account balance fits β₀ = −4 and β₁ = 0.05. For a balance of 100, the linear part is −4 + 0.05 × 100 = 1.00, and the predicted probability is 1/(1 + e^−1.00) ≈ 0.7311.

For a balance of 150, the linear part is 3.50 and the probability is ≈ 0.9707. The same 50-unit increase moved the log-odds by exactly 2.5 in both cases, but the probability rose by 0.24 in the first instance and would rise far less from an already-high starting point. That is the non-linearity the logistic function introduces.

The coefficient is best communicated as an odds ratio: e^0.05 ≈ 1.051, so each additional unit of balance multiplies the odds of default by about 1.05. Note this is a multiplicative statement about odds, not about probability.

Frequently Asked Questions

Why is a classification method called regression?

Because it regresses a continuous quantity, the log-odds, on the predictors. Classification happens afterwards, when the fitted probability is compared against a threshold. The name describes the underlying mechanism rather than the end use.

Must the classification threshold be 0.5?

No. It is a decision about the relative cost of false positives and false negatives, not a property of the model. For imbalanced problems or asymmetric costs a different threshold is often far more appropriate.

How does it compare with a linear SVM?

Both produce a linear boundary, but they optimize different objectives. Logistic regression maximizes likelihood and outputs calibrated probabilities; a support vector machine maximizes the margin and outputs a decision without a probability attached.

The Bottom Line

Logistic regression adapts the linear model to classification with two changes: a logistic squashing function so the output is a valid probability, and maximum likelihood in place of least squares. Its coefficients remain interpretable on the log-odds scale and its boundary remains linear.