Skip to content
Kudos AI

Linear Discriminant Analysis

A generative classifier that models each class as a Gaussian and inverts those models with Bayes’ theorem; assuming one covariance matrix shared by all classes gives a linear decision boundary, and one per class gives a quadratic one.

Also known as: LDA, Quadratic discriminant analysis, QDA

Understanding Linear Discriminant Analysis

Discriminant analysis approaches classification from the generative side. Instead of estimating the probability of a class given the features, it estimates how the features are distributed within each class, together with how common each class is, and then applies Bayes’ theorem to turn those into a posterior. The assumed within-class distribution is Gaussian, with a mean vector per class.

The covariance assumption is what separates the two members of the family. If every class is given the same covariance matrix, the quadratic term in the exponent is identical across classes and cancels when the posteriors are compared, so the discriminant function is linear in the features and the boundary between any two classes is a hyperplane. That is linear discriminant analysis. Allowing each class its own covariance leaves the quadratic term in place and produces curved boundaries: quadratic discriminant analysis.

The trade is parameters against flexibility. With p features and K classes, a shared covariance costs one p × (p+1)/2 matrix; separate covariances cost K of them. On 20 features that is 210 parameters against 420 for two classes, and every one of them has to be estimated from data. Quadratic discriminant analysis therefore wins when the shared-covariance assumption is genuinely wrong and the sample is large, and loses when it is right or the sample is small.

Compared with logistic regression, which produces the same shape of boundary, the difference is what is assumed and what is paid for it. Where the Gaussian assumption approximately holds, discriminant analysis is the more efficient estimator; where it does not, it is biased in a way logistic regression is not. Discriminant analysis also degrades gracefully when the classes are well separated - the situation in which logistic regression’s coefficient estimates become unstable and diverge.

How to Calculate

δₖ(x) = xᵀ Σ⁻¹ μₖ − ½ μₖᵀ Σ⁻¹ μₖ + log πₖ

where

δₖ(x)
the discriminant score for class k; the prediction is the class with the largest score
μₖ
the mean vector of class k, estimated by the class’s sample mean
Σ
the covariance matrix shared by all classes, estimated by pooling within-class deviations
πₖ
the prior probability of class k, usually the observed class proportion

Example of Linear Discriminant Analysis

Two classes in the plane, one centred at the origin with correlation +0.75 and the other centred at (1.5, 1.5) with correlation −0.75. No single covariance matrix describes both, so the optimal boundary is genuinely curved, and its error rate - computed by integrating the mixture - is 0.092708.

Fitted to 200 training points and scored on 200,000 test points, linear discriminant analysis achieves 0.118955 and quadratic discriminant analysis 0.093480. Pooling the two covariances averages a +0.75 correlation with a −0.75 one, producing something near the identity that describes neither class; the quadratic version recovers boundaries that curve and lands within 0.0008 of the floor.

The reverse case is just as instructive. On a second problem whose true boundary really is linear, averaged over 400 runs, the linear version beats the quadratic one by 0.028208 at 20 training points but by only 0.000154 at 2,000. The quadratic model is not wrong there - it contains the linear one - it simply cannot afford its own parameters until the sample is large.

Frequently Asked Questions

When should quadratic discriminant analysis be preferred?

When there is evidence that the classes have genuinely different covariance structures and the training set is large enough to estimate a separate covariance matrix per class. With few observations relative to the number of features, the shared-covariance version is usually the better bet even when it is mildly wrong.

How does it relate to naive Bayes?

Both are generative and both apply Bayes’ theorem to class-conditional densities. Naive Bayes assumes the features are independent within a class, which is a diagonal covariance matrix; discriminant analysis estimates the off-diagonal terms instead of assuming them away.

Does it require the features to be Gaussian?

The derivation does, but the method is fairly robust to moderate departures, especially in the shared-covariance form where only the pooled second moments matter. Strongly skewed or heavily discrete features are a genuine problem, and a transformation is usually worth trying first.

The Bottom Line

Discriminant analysis models the classes rather than the boundary, and the covariance assumption decides the shape it can draw. Share one covariance for a linear boundary and an efficient estimate; give each class its own for a curved boundary you will need a large sample to afford.