A Hundred Thousand Samples, Four Hundred of Them Real
On the burglary network with both neighbours calling, rejection sampling keeps 183 of 100,000 draws and likelihood weighting keeps all of them at an effective sample size of 396. Both estimates are about 10% off a posterior of 0.284172, and the reason is exactly computable: 252 samples carry 76% of the weight and 99.975% of the squared weight.
Prerequisites: Bayesian Networks and Probabilistic Inference
The burglary network is five binary variables: a burglary and an earthquake can set off an alarm, and two neighbours, John and Mary, may or may not call when it goes off. Both call. What is the probability there was a burglary?
By enumeration the answer is exact:
Now do it by sampling, with a hundred thousand samples, which sounds like more than enough for two significant figures.
A. Rejection sampling keeps 183
Draw a full assignment from the network, discard it unless both neighbours called, average the burglaries among the survivors. From 100,000 draws:
That is exactly what predicts. The evidence is rare, so almost every sample is thrown away, and the estimate rests on 183 of them. It is off by 0.027, about 10% of the quantity being estimated.
This is not an implementation flaw. The acceptance rate is the probability of the evidence, so the cost of rejection sampling scales as and any reasonably specific observation makes it hopeless. Ten binary observations at even odds would keep one sample in a thousand.
B. Likelihood weighting keeps all of them, and is worth 396
The standard fix is to stop sampling the evidence variables. Fix them to their observed values, sample everything else, and give each sample a weight equal to the probability of the evidence given its parents. Nothing is discarded.
The estimate now uses 100,000 samples, and the honest measure of how many it is worth is the effective sample size
Measured: 396.4, or 0.396% of the samples drawn, with estimate 0.2473.
So the two methods, given identical budgets, are worth 183 and 396 samples respectively. Likelihood weighting is ahead by a factor of two, and on this seed its point estimate is the further of the two from the truth, which is the right lesson about what 400 samples buys.
In the figure, switch to P(burglary | both calls). Its samples drawn slider stops at 50,000 on a fixed seed, so rejection keeps 102 samples rather than the 183 of 100,000 here, and the figure does not show the effective sample size.
Interactive: two samplers against the exact answer
The exact value comes from enumerating the joint, not from either sampler.
- Exact
- 0.3000
- Rejection sampling
- 0.3045
- Likelihood weighting
- 0.3029
- Kept by rejection
- 289
Rejection sampling reads 0.3045 from the 289 samples it kept of 1,000; likelihood weighting reads 0.3029 from all of them; the exact answer is 0.3000. Both estimators are consistent, which is a claim about a limit rather than about any one run, so the useful thing to do is drag the sample size and watch the two numbers walk toward the third. The discarded share is not noise in the procedure - it is 70.00% of the work, and it grows exponentially with the number of evidence variables.
C. Where the 396 comes from
The weight in this network takes exactly two values, because only the alarm's state matters to John and Mary:
The alarm fires with probability 0.002516, so in 100,000 samples about 251.6 of them get the large weight. Those 252 samples carry
- 76.07% of the total weight, and
- 99.975% of the total squared weight,
and squared weight is what the ESS denominator counts. The prediction is , against 396.4 measured on one run.
Nothing is wrong with the sampler. Almost every sample it draws is a world in which the alarm did not go off, and in such a world both neighbours calling is a coincidence worth a weight of 0.0005. The other 99,748 samples are not information; they are a rounding error with a variance.
D. What to do
- Report ESS, never the number of samples drawn. It is three lines and it is the only number that describes what the estimate rests on. An ESS of 396 from 100,000 draws is a result; "100,000 samples" is a budget.
- Expect ESS to fall as evidence accumulates. Every additional observation multiplies the weights by another factor below one, and in a network where the evidence is downstream of a rare cause, the collapse happens immediately rather than gradually.
- Sample from something closer to the posterior. That is what Gibbs sampling and the rest of MCMC are for: they spend their effort where the posterior actually is, instead of drawing from the prior and hoping.
- Treat a large sample count in a report as an unanswered question. It says what was spent, not what was learned.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.