Skip to content
Kudos AI

Principal Component Analysis

A technique that re-expresses data in new uncorrelated coordinates ordered by how much variance each explains, allowing dimension reduction by keeping only the first few.

Understanding Principal Component Analysis

High-dimensional data is often far less complex than its number of columns suggests, because variables move together. Principal component analysis finds the directions along which the data actually varies. The first principal component is the direction of greatest variance; the second is the direction of greatest remaining variance orthogonal to the first; and so on.

Mathematically these directions are the eigenvectors of the covariance matrix, and the corresponding eigenvalues give the variance along each. Because the eigenvectors of a symmetric matrix are orthogonal, the new coordinates are uncorrelated by construction, which is often useful in itself when the original variables are heavily collinear.

Dimension reduction follows from the ordering. Because components are sorted by variance explained, the first few frequently account for most of the total, and the remainder can be discarded with modest loss. The proportion of variance explained by each component, plotted against component number, is the standard tool for deciding how many to keep. James and colleagues describe reading an elbow off this plot, noting for one dataset that the bend after roughly the seventh component suggests little benefit in examining more.

The main cost is interpretability. A principal component is a weighted combination of every original variable, so it rarely corresponds to anything nameable. PCA also assumes the interesting structure is linear and aligned with high variance, neither of which is guaranteed: a low-variance direction can be exactly the one that separates classes, and PCA is unsupervised, so it does not know or care what the eventual prediction target is.

How to Calculate

Σ vₖ = λₖ vₖ, PVEₖ = λₖ / Σⱼ λⱼ

where

Σ
the covariance matrix of the (standardized) data
vₖ
the k-th principal component direction, an eigenvector of Σ
λₖ
the variance along that direction, its eigenvalue
PVEₖ
proportion of variance explained by component k

Example of Principal Component Analysis

Suppose height in centimetres and height in inches are both recorded. They are perfectly correlated, so the data lies on a line in two dimensions. PCA finds that the first component runs along that line and explains essentially all the variance, while the second explains almost none. One dimension has been shown to be redundant.

With a hundred correlated survey questions, the first ten components might together explain 80% of the variance. Replacing the hundred columns with those ten loses a fifth of the variability but removes a ninetieth of the columns, and downstream models fitted on the components are far less prone to overfitting.

Standardization matters before any of this. PCA maximizes variance, and variance depends on units, so a variable measured in metres and one measured in millimetres are not comparable. Without standardizing, the first component will simply follow whichever variable happens to have the largest numerical scale.

Frequently Asked Questions

How many components should be kept?

There is no exact rule. Common approaches are keeping enough components to reach a target cumulative proportion of variance explained, or looking for an elbow in the scree plot where additional components stop contributing meaningfully.

Why must variables be standardized first?

Because PCA maximizes variance, which is scale-dependent. An unstandardized variable with a large numerical range will dominate the first component purely because of its units, not because it carries more information.

Can PCA hurt predictive performance?

Yes. PCA is unsupervised and ignores the response, so it can discard a low-variance direction that happens to be highly predictive. It reduces dimension and collinearity; it does not select the features most useful for prediction.

The Bottom Line

PCA rotates data onto uncorrelated axes ordered by variance explained, letting most of the variability be retained in far fewer dimensions. It costs interpretability, assumes linear structure, and, being unsupervised, may discard exactly the direction a downstream model needed.