Bayes' Theorem and Belief Updating
Derive Bayes' theorem from the definition of conditional probability, then work the base-rate example that fools almost everyone, twice: once with the formula and once by pure counting.
Prerequisites: Probability from Zero
Probability from Zero defined the conditional probability and left one question open: what if the conditional you want is the reverse of the one you have? A diagnostic test tells you how often it fires when a condition is present. What you actually want to know is whether the condition is present given that it fired.
Bayes' theorem performs that reversal. It is two lines of algebra, and its consequences are startling enough that professionals get them wrong routinely.
A. Deriving the theorem
Recall the definition of conditional probability, written both ways round:
Both contain the same joint probability . Solving the second for it gives , and substituting into the first yields Bayes' theorem:
That is the entire derivation. No new assumption entered; it is a rearrangement of the definition. The names of the pieces matter, though:
- is the prior - what you believed before the evidence.
- is the likelihood - how the evidence behaves when holds.
- is the posterior - the updated belief.
- is the evidence, a normalising constant.
Base-rate calculator
A positive result is not the same as having the condition.
- P(condition | positive)
- 15.4%
- P(healthy | negative)
- 99.9%
- False alarms per detection
- 5.5
Out of 10,000 people tested
- Have the condition · Test positive90
- Do not have it · Test positive495
- Have the condition · Test negative10
- Do not have it · Test negative9,405
The test finds 90 genuine cases, but also flags 495 healthy people. Since a positive could be either, the chance it is real is just 15.4%. Raise the prevalence and watch it climb - the base rate, not the accuracy of the test, is doing most of the work.
B. The law of total probability
The denominator is rarely handed to you directly. You compute it by splitting the world into exhaustive, mutually exclusive cases and summing:
For a two-case split into and , that is just
This is the denominator of essentially every Bayes calculation, and recognising it on sight saves a great deal of confusion.
C. The worked example that surprises everyone
A screening test looks for a condition present in 1% of the population. The test is described as "90% accurate", which here means:
- Sensitivity: it fires for 90% of people who have the condition, so .
- False positive rate: it also fires for 5% of people who do not, so .
You test positive. What is the probability you have the condition?
Step 1 - the prior. , so .
Step 2 - the likelihoods. and , as given.
Step 3 - total probability of a positive result.
Step 4 - divide.
About 15.4%. Despite a test that sounds highly accurate, a positive result still means you probably do not have the condition.
D. Why: the base rate dominates
The reason is visible in step 3. The condition is rare, so the 5% false-positive rate is applied to a very large healthy group (99% of people) while the 90% sensitivity is applied to a tiny group (1%). The false positives - - simply outnumber the true positives - - by more than five to one.
This is the base-rate effect, and it is why a raw "accuracy" figure is close to meaningless for a rare condition without the prevalence alongside it.
E. The same answer by counting, with no formula at all
If the algebra still feels slippery, here is the identical calculation as pure bookkeeping. Imagine 10,000 people:
| Group | Count | Test positive | Test negative |
|---|---|---|---|
| Have the condition (1%) | 100 | 10 | |
| Healthy (99%) | 9,900 | 9,405 | |
| Total | 10,000 | 585 | 9,415 |
Of the 585 people who test positive, only 90 actually have the condition:
Exactly the 15.4% the formula gave, now visible as counting. Presenting Bayes problems as natural frequencies like this makes them dramatically easier to reason about, and is worth doing whenever you need to explain a result to someone else.
F. Checking it in code
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Both routes print , confirming the table and the formula agree.
G. Updating twice: yesterday's posterior is today's prior
Bayes composes. Take a second, independent test that also comes back positive. The posterior from the first test, , becomes the new prior:
Two positives push the belief from 15.4% to about 76.6%. The evidence was never weak; it was fighting an extremely low base rate, and one round was not enough to overcome it.
The independence caveat, again. Chaining updates like this assumes the two tests fail independently given the true state. If they share a failure mode - > the same reagent, the same miscalibrated instrument - a second positive carries far less information than the calculation credits it with, and 76.6% is an overstatement.
Key takeaways
- Bayes' theorem is a rearrangement of the definition of conditional probability, not an extra assumption.
- The four pieces are prior, likelihood, posterior, and the evidence denominator, which you usually compute with the law of total probability.
- With a rare condition, false positives can swamp true positives: a "90% accurate" test gave a positive predictive value of only 15.4%.
- Re-expressing the problem in natural frequencies (90 out of 585) gives the identical answer with no algebra.
- Updates compose - the posterior becomes the next prior - but only if the pieces of evidence are conditionally independent.
What's next
We have been handed probabilities so far. The rest of this site is mostly about the opposite problem: estimating them from data, and knowing how much to trust the estimate. That is the subject of What Is Statistical Learning?, which introduces the split between the error you can reduce and the error you cannot.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.