Skip to content
Kudos AI
Lire en français
Probability Foundations

Which Wrong Distribution Do You Want?

One bimodal target, one Gaussian, and two directions of the same divergence. Minimising KL(P||Q) puts the Gaussian across both modes with almost no mass where the target actually lives; minimising KL(Q||P) puts it on one mode at a value of 0.6931 nats, which is ln 2 to four decimals and not a coincidence. Each fit is judged catastrophic by the other objective, 2.0976 against 15.2799.

4 min readKudos AI

Prerequisites: Entropy and Information

A two-humped target with a single bell sliding across it, first stretched to span both humps and then narrowed onto one, with the divergence read off in both directions.

The target PP is an even mixture of two unit Gaussians at −4-4 and +4+4. It has mean 00, standard deviation 4.12314.1231, and essentially no probability anywhere near its own mean.

The model QQ is a single Gaussian. It cannot represent PP, which is the normal situation: the question is what "as close as possible" means.

The Kullback-Leibler divergence gives two answers, because it is not symmetric.

A. Two fits

D(P∥Q)=∫plog⁡pq,D(Q∥P)=∫qlog⁡qp.D(P \parallel Q) = \int p \log \frac{p}{q}, \qquad D(Q \parallel P) = \int q \log \frac{q}{p}.

Minimise each over the mean and standard deviation of QQ:

Objectivefitted meanfitted sdvalue
D(P∥Q)D(P \parallel Q)004.12314.12310.7236 nats
D(Q∥P)D(Q \parallel P)+3.9996+3.99961.00081.00080.6931 nats

The first fit spans both modes. Its standard deviation, 4.1231, is exactly the standard deviation of PP itself, 1+16=17\sqrt{1 + 16} = \sqrt{17}: minimising D(P∥Q)D(P \parallel Q) over Gaussians matches the mean and variance of the target, and here that puts the centre at the point where PP has almost no mass at all. The second fit lands on one mode, within a thousandth: mean 3.9996 against a true mode at 4, standard deviation 1.0008 against a true 1. The small offsets are the other component's tail, reaching faintly across the gap.

Both answers are correct. They are answers to different questions.

Interactive: which way round the divergence is taken

Both divergences in nats, integrated over the whole line.

-8-4048P, the targetQ, one Gaussian
KL(P || Q), mean-seeking
0.7236
KL(Q || P), mode-seeking
2.0976

This is the forward optimum, and it is moment matching exactly: Q takes P’s own mean, 0, and standard deviation, the square root of 17, 4.1231. So it spans both modes and is centred where P has almost no mass. Forward KL scores it 0.7236; reverse KL, which charges Q for every stretch it puts where P is empty, scores the same Q 2.0976.

B. Why 0.6931

The reverse-KL value is ln⁡2=0.6931\ln 2 = 0.6931 to four decimals, and that is not a coincidence. When QQ sits on one component, P≈12QP \approx \tfrac{1}{2}Q throughout the region where QQ has mass, since the other component contributes almost nothing there. So

D(Q∥P)≈∫qlog⁡q12q=log⁡2.D(Q \parallel P) \approx \int q \log \frac{q}{\tfrac{1}{2} q} = \log 2.

The "almost" can be made exact. With Q=N(4,1)Q = N(4, 1), the target is p=12q (1+e−8x)p = \tfrac{1}{2} q \,(1 + e^{-8x}), so

D(Q∥P)=ln⁡2−Eq ⁣[ln⁡(1+e−8x)],D(Q \parallel P) = \ln 2 - \mathbb{E}_q\!\left[\ln\left(1 + e^{-8x}\right)\right],

and the correction, the other component's tail reaching across the gap, is about 0.00010.0001. The fitted value is 0.6930530.693053 against ln⁡2=0.693147\ln 2 = 0.693147.

The penalty for ignoring half the target is, up to that sliver, the one bit of information needed to say which half was kept. That is a clean way to see what reverse KL does and does not charge for: it charges for putting mass where the target has none, and charges nothing at all for failing to cover the target's other regions.

C. Each fit is a disaster by the other measure

judged by D(P∥Q)D(P \parallel Q)judged by D(Q∥P)D(Q \parallel P)
the forward fit0.72362.0976
the reverse fit15.27990.6931

The reverse fit scores 15.2799 under forward KL, twenty-one times worse than the forward fit's own 0.7236. This is not a small preference between two nearly equivalent criteria; the two objectives disagree about which model is usable.

The asymmetry has a direction that is easy to remember:

  • Forward, D(P∥Q)D(P \parallel Q), is mean-seeking. The integral is weighted by pp, so wherever PP has mass and QQ does not, the ratio p/qp/q blows up. QQ is forced to cover everything PP does, at the cost of covering a great deal that PP does not.
  • Reverse, D(Q∥P)D(Q \parallel P), is mode-seeking. The integral is weighted by qq, so regions where PP has mass and QQ does not cost nothing. QQ is free to ignore most of PP, provided it avoids putting mass where PP has none.

D. Which one you are already using

Both are in daily use, usually without the direction being stated.

  • Maximum likelihood is forward KL. Minimising cross-entropy against the data distribution minimises D(P∥Q)D(P \parallel Q) up to a constant. That is why a maximum-likelihood fit spreads: it is being penalised for every region of the data it fails to cover.
  • Variational inference is reverse KL. The evidence lower bound minimises D(Q∥P)D(Q \parallel P), which is why a variational posterior is famously overconfident and latches onto one mode.
  • Distillation, RLHF and their relatives choose a direction too, and the choice shows up as whether the student hedges across the teacher's options or commits to one of them.

The practical rule: if leaving out part of the target is the expensive mistake, use forward. If claiming things the target rules out is the expensive mistake, use reverse. And when someone reports "the KL divergence" as one number, the first question is which way round it was taken.

References & further reading

  • David J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge University Press, 2003source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readInformation Theory

The Bound That Is Actually Reached

Entropy is not a summary of a distribution but a floor that the best code meets to the last decimal, the surcharge for using the wrong distribution is exactly the loss every classifier already minimises, and mutual information puts a hard ceiling on everything downstream of a sensor. Three results, each unusually sharp.

MathematicsMachine Learning
10 min readProbability Foundations

Entropy and Information

Measuring uncertainty in bits: Shannon entropy and why the logarithm is base 2, information gain worked on a split, and how cross-entropy and KL divergence relate to entropy and to the loss functions used to train classifiers.

Information TheoryProbabilityMathematics
3 min readProbability Foundations

The Two Features That Look Like Noise

A variable that determines another with a correlation of exactly 0.0000000000, and a pair of features whose every pairwise mutual information with the target is exactly zero while the two together determine it completely. Univariate screening discards both, and the second case is the one that matters: the features it removes are removed because they matter.

Machine LearningMathematics
← Back to all articles