Separating Hyperplanes and the Margin
A hyperplane as a decision rule, the margin as the width of the widest slab between the classes, the handful of observations that fix it, and the two ways the idea fails.
Logistic regression fits a boundary by making the observed labels probable. This lesson takes a different starting point: among all boundaries that separate the classes, prefer the one that is furthest from every observation. That single idea produces a classifier with an unusual property, which is that almost all of the data turn out to be irrelevant to it.
A hyperplane is a decision rule
In dimensions a hyperplane is the flat, -dimensional set
In two dimensions that is a line, in three a plane. A point not on it makes the left-hand side either positive or negative, so a hyperplane cuts the space in two and gives a classifier for free: write
and predict class when and class when . Coding the labels as makes "correctly classified" compact: the prediction is right exactly when
The magnitude of is informative too. A point far from the boundary has far from zero, and we can be confident about it; a point close to the boundary is close to a coin flip.
Which separating hyperplane?
If the classes can be separated at all, they can usually be separated in infinitely many ways: nudge or tilt a separating line slightly and it still separates. So "find a separating hyperplane" is not yet a well-posed problem.
The maximal margin classifier resolves it. Compute the distance from every training observation to a candidate hyperplane; the smallest of those distances is the margin. Then choose the hyperplane whose margin is largest. In words, it is the mid-line of the widest slab you can push between the two classes.
Worked example
Take six observations in the plane:
The maximal margin hyperplane is
and the perpendicular distance from a point to it is . Evaluating that at each observation:
The margin is , and four observations achieve it: , , and . These are the support vectors. They lie on the edges of the slab and they hold it in place - move one and the hyperplane moves.
The other two, and , sit at and are irrelevant. You may move them anywhere on their side of the slab and the fitted classifier does not change at all. Note also that the support vectors are not evenly split between the classes: one positive, three negative.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Interactive: the widest slab that fits
Ringed points are the support vectors.
- Margin
- 1.414214
- Support vectors
- 4
- Separable
- yes
The widest slab has half-width 1.414214, and exactly 4 points touch it. Those are the support vectors, and they are the whole solution: the two positives further out sit at 2.8284 and could be moved anywhere on their own side, outside the slab, without the boundary twitching. A classifier that depends on four points out of six is a strange object, and it is the reason margins generalise well and are fragile at the same time.
Stating it as an optimisation
The problem is written
The second constraint says every observation is on its correct side and at least away. The first looks like a technicality and is not. Multiplying every coefficient by any describes the same hyperplane, so without a scale convention the parameters are not determined. Fixing the coefficient vector to unit length pins them down, and it makes equal the actual perpendicular distance - which is what makes "maximise " mean "maximise the margin".
The canonical form. An equivalent convention scales so the closest points satisfy . Our hyperplane becomes with , giving and a margin of , the same . Maximising the margin is then minimising , which is the form most software solves.
Two ways it fails
The classes may not be separable. Then no hyperplane satisfies the constraints with and the problem simply has no solution. Real data are frequently like this, and a method that returns nothing at all is not much use.
Even when it works, it is fragile. The solution is determined entirely by the few points nearest the boundary, so it inherits their instability. Add a single observation at , tucked in near the negative cloud, and the best achievable margin falls from to - a factor of more than three, from one point. Since a narrow margin is precisely what generalises poorly, demanding perfect separation is self-defeating.
Both failures point the same way: the classifier should be allowed to get a few observations wrong. That is the next lesson.
Before the quiz
Be able to classify by the sign of , define the margin, pick out the support vectors and say why the others do not matter, explain what the normalisation constraint is for, and name the two failure modes.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.
Unlock the full path
This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.