Skip to content
Kudos AI

Bayes’ Theorem

A rule for updating the probability of a hypothesis in light of new evidence, by inverting a conditional probability.

Also known as: Bayes’ rule, Bayes’ law

A highly accurate test for a rare condition, with the false positives outnumbering the true ones on screen.

Understanding Bayes’ Theorem

Many questions have a natural direction that is the reverse of what can be measured. A diagnostic test is characterized by how often it fires when the disease is present, that is P(positive | disease). The clinically useful quantity is the opposite conditional, P(disease | positive). Bayes’ theorem is the identity that connects them.

It follows directly from the definition of conditional probability. Since P(A ∩ B) can be written either as P(A | B)P(B) or as P(B | A)P(A), equating the two and dividing rearranges into the familiar form. The theorem is therefore not an assumption or a modelling choice; it is an algebraic consequence of the probability axioms.

The interpretive content sits in the three quantities on the right. The prior P(A) is what was believed before the evidence arrived. The likelihood P(B | A) says how expected the evidence would be if the hypothesis were true. The normalizing term P(B) is the total probability of seeing that evidence at all, across every hypothesis. Their combination gives the posterior P(A | B).

The dominant practical failure is neglecting the prior. When a condition is rare, even an accurate test produces mostly false positives, because the large population of healthy people generates more false alarms than the small population of sick people generates true ones. Reasoning from the test’s accuracy alone, without the base rate, gives an answer that can be wrong by an order of magnitude.

How to Calculate

P(A | B) = P(B | A) · P(A) / P(B)

where

P(A | B)
the posterior: probability of A after observing B
P(B | A)
the likelihood: probability of observing B if A holds
P(A)
the prior: probability of A before observing B
P(B)
the evidence, or marginal likelihood, of observing B at all

Example of Bayes’ Theorem

Suppose a disease affects 1 person in 1,000. A test detects it in 99% of those who have it, and returns a false positive for 5% of those who do not. Someone tests positive. What is the probability they are ill?

Take 100,000 people. About 100 have the disease, and the test correctly flags 99 of them. The other 99,900 are healthy, and 5% of them, about 4,995, are flagged anyway. So 99 + 4,995 = 5,094 positive results occur in total, of which only 99 are genuine.

The probability of disease given a positive test is therefore 99 / 5,094 ≈ 0.0194, under 2%. The test is genuinely accurate, yet the great majority of its positives are false, purely because the healthy population is a thousand times larger. This is the base-rate effect, and it is why a positive screening result is normally followed by a second, independent test.

Frequently Asked Questions

What is the difference between the prior and the posterior?

The prior is the probability assigned to a hypothesis before the evidence is taken into account; the posterior is the revised probability afterwards. Applying Bayes’ theorem repeatedly, the posterior from one round of evidence becomes the prior for the next.

Where does the prior come from?

From whatever is known beforehand: a known base rate, historical frequency, or a deliberately uninformative choice when little is known. The dependence on a prior is the main point of contention between Bayesian and frequentist approaches, though when data is plentiful the posterior is dominated by the likelihood and the choice of prior matters less.

Why is the denominator often ignored in practice?

P(B) does not depend on the hypothesis, so when comparing hypotheses against the same evidence it is a common constant. It can be dropped when only the relative ranking matters, which is exactly what a naive Bayes classifier does when it picks the most probable class.

The Bottom Line

Bayes’ theorem is the rule for turning evidence into a revised belief, and its practical force lies in insisting that the prior be included. Skipping the base rate is what makes accurate tests look far more conclusive than they are.