Skip to content
Kudos AI
Lire en français
Probability Foundations

Bayes' Theorem and Belief Updating

Derive Bayes' theorem from the definition of conditional probability, then work the base-rate example that fools almost everyone, twice: once with the formula and once by pure counting.

6 min readKudos AI

Prerequisites: Probability from Zero

Ten thousand people, split into the 100 who are ill and the 9,900 who are not, then thinned to who tests positive: 90 real cases against 495 false alarms, which is why a 90%-accurate test gives a posterior of 15.4%.

Probability from Zero defined the conditional probability P(a∣b)P(a \mid b) and left one question open: what if the conditional you want is the reverse of the one you have? A diagnostic test tells you how often it fires when a condition is present. What you actually want to know is whether the condition is present given that it fired.

Bayes' theorem performs that reversal. It is two lines of algebra, and its consequences are startling enough that professionals get them wrong routinely.

A. Deriving the theorem

Recall the definition of conditional probability, written both ways round:

P(a∣b)=P(a and b)P(b),P(b∣a)=P(a and b)P(a).P(a \mid b) = \frac{P(a \text{ and } b)}{P(b)}, \qquad P(b \mid a) = \frac{P(a \text{ and } b)}{P(a)} .

Both contain the same joint probability P(a and b)P(a \text{ and } b). Solving the second for it gives P(a and b)=P(b∣a) P(a)P(a \text{ and } b) = P(b \mid a)\,P(a), and substituting into the first yields Bayes' theorem:

P(a∣b)=P(b∣a) P(a)P(b).P(a \mid b) = \frac{P(b \mid a)\, P(a)}{P(b)} .

That is the entire derivation. No new assumption entered; it is a rearrangement of the definition. The names of the pieces matter, though:

  • P(a)P(a) is the prior - what you believed before the evidence.
  • P(b∣a)P(b \mid a) is the likelihood - how the evidence behaves when aa holds.
  • P(a∣b)P(a \mid b) is the posterior - the updated belief.
  • P(b)P(b) is the evidence, a normalising constant.

Base-rate calculator

A positive result is not the same as having the condition.

P(condition | positive)
15.4%
P(healthy | negative)
99.9%
False alarms per detection
5.5

The test finds 90 genuine cases, but also flags 495 healthy people. Since a positive could be either, the chance it is real is just 15.4%. Raise the prevalence and watch it climb - the base rate, not the accuracy of the test, is doing most of the work.

B. The law of total probability

The denominator P(b)P(b) is rarely handed to you directly. You compute it by splitting the world into exhaustive, mutually exclusive cases and summing:

P(b)=∑iP(b∣ai) P(ai).P(b) = \sum_i P(b \mid a_i)\, P(a_i) .

For a two-case split into aa and ¬a\neg a, that is just

P(b)=P(b∣a)P(a)+P(b∣¬a)P(¬a).P(b) = P(b \mid a)P(a) + P(b \mid \neg a)P(\neg a) .

This is the denominator of essentially every Bayes calculation, and recognising it on sight saves a great deal of confusion.

C. The worked example that surprises everyone

A screening test looks for a condition present in 1% of the population. The test is described as "90% accurate", which here means:

  • Sensitivity: it fires for 90% of people who have the condition, so P(positive∣condition)=0.90P(\text{positive} \mid \text{condition}) = 0.90.
  • False positive rate: it also fires for 5% of people who do not, so P(positive∣healthy)=0.05P(\text{positive} \mid \text{healthy}) = 0.05.

You test positive. What is the probability you have the condition?

Step 1 - the prior. P(condition)=0.01P(\text{condition}) = 0.01, so P(healthy)=0.99P(\text{healthy}) = 0.99.

Step 2 - the likelihoods. 0.900.90 and 0.050.05, as given.

Step 3 - total probability of a positive result.

P(positive)=0.90×0.01⏟true positives+0.05×0.99⏟false positives=0.009+0.0495=0.0585.P(\text{positive}) = \underbrace{0.90 \times 0.01}_{\text{true positives}} + \underbrace{0.05 \times 0.99}_{\text{false positives}} = 0.009 + 0.0495 = 0.0585 .

Step 4 - divide.

P(condition∣positive)=0.90×0.010.0585=0.0090.0585≈0.1538.P(\text{condition} \mid \text{positive}) = \frac{0.90 \times 0.01}{0.0585} = \frac{0.009}{0.0585} \approx 0.1538 .

About 15.4%. Despite a test that sounds highly accurate, a positive result still means you probably do not have the condition.

D. Why: the base rate dominates

The reason is visible in step 3. The condition is rare, so the 5% false-positive rate is applied to a very large healthy group (99% of people) while the 90% sensitivity is applied to a tiny group (1%). The false positives - 0.04950.0495 - simply outnumber the true positives - 0.0090.009 - by more than five to one.

This is the base-rate effect, and it is why a raw "accuracy" figure is close to meaningless for a rare condition without the prevalence alongside it.

E. The same answer by counting, with no formula at all

If the algebra still feels slippery, here is the identical calculation as pure bookkeeping. Imagine 10,000 people:

GroupCountTest positiveTest negative
Have the condition (1%)100100×0.90=90100 \times 0.90 = 9010
Healthy (99%)9,9009,900×0.05=4959{,}900 \times 0.05 = 4959,405
Total10,0005859,415

Of the 585 people who test positive, only 90 actually have the condition:

90585≈0.1538.\frac{90}{585} \approx 0.1538 .

Exactly the 15.4% the formula gave, now visible as counting. Presenting Bayes problems as natural frequencies like this makes them dramatically easier to reason about, and is worth doing whenever you need to explain a result to someone else.

F. Checking it in code

Python

Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.

Both routes print 0.15380.1538, confirming the table and the formula agree.

G. Updating twice: yesterday's posterior is today's prior

Bayes composes. Take a second, independent test that also comes back positive. The posterior from the first test, 0.15380.1538, becomes the new prior:

P(positive2)=0.90×0.1538+0.05×0.8462=0.1384+0.0423=0.1807,P(\text{positive}_2) = 0.90 \times 0.1538 + 0.05 \times 0.8462 = 0.1384 + 0.0423 = 0.1807 , P(condition∣both positive)=0.90×0.15380.1807≈0.766.P(\text{condition} \mid \text{both positive}) = \frac{0.90 \times 0.1538}{0.1807} \approx 0.766 .

Two positives push the belief from 15.4% to about 76.6%. The evidence was never weak; it was fighting an extremely low base rate, and one round was not enough to overcome it.

The independence caveat, again. Chaining updates like this assumes the two tests fail independently given the true state. If they share a failure mode - > the same reagent, the same miscalibrated instrument - a second positive carries far less information than the calculation credits it with, and 76.6% is an overstatement.

Key takeaways

  • Bayes' theorem is a rearrangement of the definition of conditional probability, not an extra assumption.
  • The four pieces are prior, likelihood, posterior, and the evidence denominator, which you usually compute with the law of total probability.
  • With a rare condition, false positives can swamp true positives: a "90% accurate" test gave a positive predictive value of only 15.4%.
  • Re-expressing the problem in natural frequencies (90 out of 585) gives the identical answer with no algebra.
  • Updates compose - the posterior becomes the next prior - but only if the pieces of evidence are conditionally independent.

What's next

We have been handed probabilities so far. The rest of this site is mostly about the opposite problem: estimating them from data, and knowing how much to trust the estimate. That is the subject of What Is Statistical Learning?, which introduces the split between the error you can reduce and the error you cannot.

References & further reading

  • Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readProbability Foundations

Probability from Zero: The Language of Uncertainty

Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.

ProbabilityMathematicsArtificial Intelligence
3 min readProbabilistic Reasoning

A Hundred Thousand Samples, Four Hundred of Them Real

On the burglary network with both neighbours calling, rejection sampling keeps 183 of 100,000 draws and likelihood weighting keeps all of them at an effective sample size of 396. Both estimates are about 10% off a posterior of 0.284172, and the reason is exactly computable: 252 samples carry 76% of the weight and 99.975% of the squared weight.

Artificial IntelligenceProbability
7 min readStatistical Learning Foundations

What Is Statistical Learning?

The setup behind every predictive model: estimating an unknown function f from data, the split between reducible and irreducible error, and why prediction and inference pull in different directions.

StatisticsMachine LearningMathematics
← Back to all articles