Which Wrong Distribution Do You Want?
One bimodal target, one Gaussian, and two directions of the same divergence. Minimising KL(P||Q) puts the Gaussian across both modes with almost no mass where the target actually lives; minimising KL(Q||P) puts it on one mode at a value of 0.6931 nats, which is ln 2 to four decimals and not a coincidence. Each fit is judged catastrophic by the other objective, 2.0976 against 15.2799.
Prerequisites: Entropy and Information
The target is an even mixture of two unit Gaussians at and . It has mean , standard deviation , and essentially no probability anywhere near its own mean.
The model is a single Gaussian. It cannot represent , which is the normal situation: the question is what "as close as possible" means.
The Kullback-Leibler divergence gives two answers, because it is not symmetric.
A. Two fits
Minimise each over the mean and standard deviation of :
| Objective | fitted mean | fitted sd | value |
|---|---|---|---|
| 0.7236 nats | |||
| 0.6931 nats |
The first fit spans both modes. Its standard deviation, 4.1231, is exactly the standard deviation of itself, : minimising over Gaussians matches the mean and variance of the target, and here that puts the centre at the point where has almost no mass at all. The second fit lands on one mode, within a thousandth: mean 3.9996 against a true mode at 4, standard deviation 1.0008 against a true 1. The small offsets are the other component's tail, reaching faintly across the gap.
Both answers are correct. They are answers to different questions.
Interactive: which way round the divergence is taken
Both divergences in nats, integrated over the whole line.
- KL(P || Q), mean-seeking
- 0.7236
- KL(Q || P), mode-seeking
- 2.0976
This is the forward optimum, and it is moment matching exactly: Q takes P’s own mean, 0, and standard deviation, the square root of 17, 4.1231. So it spans both modes and is centred where P has almost no mass. Forward KL scores it 0.7236; reverse KL, which charges Q for every stretch it puts where P is empty, scores the same Q 2.0976.
B. Why 0.6931
The reverse-KL value is to four decimals, and that is not a coincidence. When sits on one component, throughout the region where has mass, since the other component contributes almost nothing there. So
The "almost" can be made exact. With , the target is , so
and the correction, the other component's tail reaching across the gap, is about . The fitted value is against .
The penalty for ignoring half the target is, up to that sliver, the one bit of information needed to say which half was kept. That is a clean way to see what reverse KL does and does not charge for: it charges for putting mass where the target has none, and charges nothing at all for failing to cover the target's other regions.
C. Each fit is a disaster by the other measure
| judged by | judged by | |
|---|---|---|
| the forward fit | 0.7236 | 2.0976 |
| the reverse fit | 15.2799 | 0.6931 |
The reverse fit scores 15.2799 under forward KL, twenty-one times worse than the forward fit's own 0.7236. This is not a small preference between two nearly equivalent criteria; the two objectives disagree about which model is usable.
The asymmetry has a direction that is easy to remember:
- Forward, , is mean-seeking. The integral is weighted by , so wherever has mass and does not, the ratio blows up. is forced to cover everything does, at the cost of covering a great deal that does not.
- Reverse, , is mode-seeking. The integral is weighted by , so regions where has mass and does not cost nothing. is free to ignore most of , provided it avoids putting mass where has none.
D. Which one you are already using
Both are in daily use, usually without the direction being stated.
- Maximum likelihood is forward KL. Minimising cross-entropy against the data distribution minimises up to a constant. That is why a maximum-likelihood fit spreads: it is being penalised for every region of the data it fails to cover.
- Variational inference is reverse KL. The evidence lower bound minimises , which is why a variational posterior is famously overconfident and latches onto one mode.
- Distillation, RLHF and their relatives choose a direction too, and the choice shows up as whether the student hedges across the teacher's options or commits to one of them.
The practical rule: if leaving out part of the target is the expensive mistake, use forward. If claiming things the target rules out is the expensive mistake, use reverse. And when someone reports "the KL divergence" as one number, the first question is which way round it was taken.
References & further reading
- David J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge University Press, 2003source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.