Skip to content
Kudos AI
Lire en français
Probabilistic Reasoning

A Hundred Thousand Samples, Four Hundred of Them Real

On the burglary network with both neighbours calling, rejection sampling keeps 183 of 100,000 draws and likelihood weighting keeps all of them at an effective sample size of 396. Both estimates are about 10% off a posterior of 0.284172, and the reason is exactly computable: 252 samples carry 76% of the weight and 99.975% of the squared weight.

3 min readKudos AI

Prerequisites: Bayesian Networks and Probabilistic Inference

Samples pouring through a network and almost all of them discarded at the evidence node, then the same samples kept with weights, most of them so light they stack to nothing.

The burglary network is five binary variables: a burglary and an earthquake can set off an alarm, and two neighbours, John and Mary, may or may not call when it goes off. Both call. What is the probability there was a burglary?

By enumeration the answer is exact:

P(Burglary∣j,m)=0.284172,P(j,m)=0.00208410.P(\text{Burglary} \mid j, m) = 0.284172, \qquad P(j, m) = 0.00208410.

Now do it by sampling, with a hundred thousand samples, which sounds like more than enough for two significant figures.

A. Rejection sampling keeps 183

Draw a full assignment from the network, discard it unless both neighbours called, average the burglaries among the survivors. From 100,000 draws:

183 kept (0.183%),P^=0.2568.183 \text{ kept } (0.183\%), \qquad \hat{P} = 0.2568.

That is exactly what P(j,m)=0.00208P(j, m) = 0.00208 predicts. The evidence is rare, so almost every sample is thrown away, and the estimate rests on 183 of them. It is off by 0.027, about 10% of the quantity being estimated.

This is not an implementation flaw. The acceptance rate is the probability of the evidence, so the cost of rejection sampling scales as 1/P(e)1/P(e) and any reasonably specific observation makes it hopeless. Ten binary observations at even odds would keep one sample in a thousand.

B. Likelihood weighting keeps all of them, and is worth 396

The standard fix is to stop sampling the evidence variables. Fix them to their observed values, sample everything else, and give each sample a weight equal to the probability of the evidence given its parents. Nothing is discarded.

The estimate now uses 100,000 samples, and the honest measure of how many it is worth is the effective sample size

ESS=(∑iwi)2∑iwi2.\text{ESS} = \frac{\left(\sum_i w_i\right)^2}{\sum_i w_i^2}.

Measured: 396.4, or 0.396% of the samples drawn, with estimate 0.2473.

So the two methods, given identical budgets, are worth 183 and 396 samples respectively. Likelihood weighting is ahead by a factor of two, and on this seed its point estimate is the further of the two from the truth, which is the right lesson about what 400 samples buys.

In the figure, switch to P(burglary | both calls). Its samples drawn slider stops at 50,000 on a fixed seed, so rejection keeps 102 samples rather than the 183 of 100,000 here, and the figure does not show the effective sample size.

Interactive: two samplers against the exact answer

The exact value comes from enumerating the joint, not from either sampler.

0.3000
Exact
0.3000
Rejection sampling
0.3045
Likelihood weighting
0.3029
Kept by rejection
289

Rejection sampling reads 0.3045 from the 289 samples it kept of 1,000; likelihood weighting reads 0.3029 from all of them; the exact answer is 0.3000. Both estimators are consistent, which is a claim about a limit rather than about any one run, so the useful thing to do is drag the sample size and watch the two numbers walk toward the third. The discarded share is not noise in the procedure - it is 70.00% of the work, and it grows exponentially with the number of evidence variables.

C. Where the 396 comes from

The weight in this network takes exactly two values, because only the alarm's state matters to John and Mary:

w=0.9×0.7=0.63if the alarm fired,w=0.05×0.01=0.0005if it did not.w = 0.9 \times 0.7 = 0.63 \quad \text{if the alarm fired}, \qquad w = 0.05 \times 0.01 = 0.0005 \quad \text{if it did not}.

The alarm fires with probability 0.002516, so in 100,000 samples about 251.6 of them get the large weight. Those 252 samples carry

  • 76.07% of the total weight, and
  • 99.975% of the total squared weight,

and squared weight is what the ESS denominator counts. The prediction is ESS=434.8\text{ESS} = 434.8, against 396.4 measured on one run.

Nothing is wrong with the sampler. Almost every sample it draws is a world in which the alarm did not go off, and in such a world both neighbours calling is a coincidence worth a weight of 0.0005. The other 99,748 samples are not information; they are a rounding error with a variance.

D. What to do

  • Report ESS, never the number of samples drawn. It is three lines and it is the only number that describes what the estimate rests on. An ESS of 396 from 100,000 draws is a result; "100,000 samples" is a budget.
  • Expect ESS to fall as evidence accumulates. Every additional observation multiplies the weights by another factor below one, and in a network where the evidence is downstream of a rare cause, the collapse happens immediately rather than gradually.
  • Sample from something closer to the posterior. That is what Gibbs sampling and the rest of MCMC are for: they spend their effort where the posterior actually is, instead of drawing from the prior and hoping.
  • Treat a large sample count in a report as an unanswered question. It says what was spent, not what was learned.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

9 min readProbabilistic Reasoning

Bayesian Networks and Probabilistic Inference

How a graph and a few small tables stand in for a joint distribution with thousands of entries, how to answer a query against it exactly by enumeration and variable elimination, and what to do when exact inference is out of reach: rejection sampling, likelihood weighting, and Gibbs sampling, each worked through on the same two networks.

ProbabilityArtificial Intelligence
5 min readProbabilistic Reasoning

The Week That Cannot Have Happened

Take the most likely state on each day and write them down in order, and you have a report the model assigns probability exactly zero: on a four-day machine-monitoring example the day-by-day answer is healthy, healthy, failed, failed, and healthy to failed is a transition that cannot occur. What the two questions actually are, why smoothing and Viterbi answer different ones, and what the 0.411 posterior on the best path means for anyone who has to act on it.

Artificial IntelligenceProbability
10 min readStatistical Inference

What a Sample Can and Cannot Tell You

Estimators as random variables with distributions of their own, the case where the unbiased estimator is the worse one, what a confidence interval actually promises and the standard interval that delivers 87% where it advertises 95%, and what a p-value is a probability of - every figure computed exactly or by fixed-seed simulation.

StatisticsProbability
← Back to all articles