Principal Components and Variance
The direction of maximum variance, the constraint that makes the question well posed, the proportion of variance explained, and why scaling is not optional.
Six points, two measurements each. By the end of this lesson one number will replace both of them and keep 98.83% of everything that varied.
That is the whole idea of principal components: most datasets with many columns are not really that wide. The columns move together, and the movement lies close to a lower-dimensional shape. Finding that shape lets you carry almost all of the information in far fewer numbers.
The catch is that there is nothing to check your answer against. Everything in the supervised path had a response to predict, so a model could be right or wrong. Here there is no response at all - only the data - and "what structure does it contain" is a genuinely harder question even to pose. So the first job is to say precisely what a good answer would mean.
Principal component analysis makes the question precise in one way: find a small number of directions along which the data vary as much as possible, and use them in place of the original variables.
The first principal component
The first principal component is the normalised linear combination
that has the largest variance. The constraint is not decoration. Without it you could double every , quadruple the variance, and repeat forever - the maximisation would have no solution at all. Fixing the length to one makes the problem about direction rather than magnitude.
The vector is the loading vector. It describes a direction in feature space, and it belongs to the dataset as a whole. Projecting the observations onto it gives the scores , one per observation.
Loadings and scores answer different questions. A loading says how much a variable contributes to a component; a score says where an observation sits along it. Confusing the two is the most common way to misread a PCA output.
James et al. add a second, equivalent characterisation worth keeping: the first component direction is also the line closest to all observations, in the sense of squared perpendicular distance. Maximising the variance of the projections and minimising the distance to the line are the same problem.
Worked example
Six observations on two variables:
Centre the columns and the sample covariance matrix is
Its eigenvalues are and , and the first eigenvector - the first loading vector - is
Both loadings are positive and the second is larger, so the component is essentially "both variables together, weighted toward the second". The scores are
and their sample variance is - exactly . That is not a coincidence: the eigenvalue is the variance captured by its component.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Interactive: what a projection keeps, and what it loses
Six points, one line, every angle available.
- Variance kept
- 2.0000
- Squared distance lost
- 6.1667
- Their sum
- 8.1667
Keeping 2.0000 of the variance. Watch the third figure as you turn the line: the first two move in opposite directions and their sum never changes, whatever angle you choose. Maximising the variance you keep and minimising the perpendicular distance to the line are not two criteria that happen to agree - they are one criterion written two ways. No angle beats the principal component; try to find one.
Proportion of variance explained
Each component captures some share of the total variance, and the proportion of variance explained (PVE) reports it:
Here and . One direction carries 98.83% of the variation in a two-variable dataset, so replacing both variables with a single score loses barely anything.
A scree plot puts the PVE of each component in order, and the usual advice is to look for an elbow. It is worth being honest about what that is: an eyeball judgement, not a test. There is no widely accepted objective way to decide how many components are enough, and in an unsupervised setting there is no held-out error to appeal to either - which is exactly the difficulty of having no .
Why standardising is not optional
PCA maximises variance, and variance has units. Measure a length in millimetres rather than metres and its variance grows by a factor of , so that variable will dominate the first component for no reason but the choice of ruler.
The fix is to scale each variable to standard deviation one before computing the components - which amounts to running PCA on the correlation matrix rather than the covariance matrix. James et al. describe this as the usual practice, and the exception is worth naming: when the variables are already in the same units and their differing variances are genuinely meaningful, scaling would throw away real information.
Before the quiz
Be able to state what the first component maximises and why the norm constraint is needed, distinguish a loading from a score, compute a PVE from eigenvalues, and say what goes wrong if you skip standardisation. The encyclopedia entry Principal Component Analysis covers the same material as a reference.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.