Probability from Zero: The Language of Uncertainty
Build probability from the ground up: possible worlds, the sample space, the two basic axioms, and the addition and multiplication rules, each derived rather than asserted, with worked numeric examples.
Almost every idea on this site eventually rests on a single sentence: we do not know what will happen, but we can say how likely each outcome is. This article makes that sentence precise. We build probability from two axioms, derive the familiar rules rather than assuming them, and finish with the tools you need for every later article in this track.
Nothing here is taken on faith. Each rule below is derived from the two axioms, so by the end you will know not just what the formulas say but why they must be true.
A. Possible worlds and the sample space
Russell & Norvig frame probability in terms of possible worlds. A probabilistic assertion says how likely each world is, where a logical assertion would say only which worlds are ruled out.
The set of all possible worlds is the sample space, written (uppercase omega). An individual world is written . Two properties define it:
- The worlds are mutually exclusive - two cannot both be the case.
- The worlds are exhaustive - one of them must be the case.
For a single roll of an ordinary die, the sample space is
For a roll of two distinguishable dice there are 36 possible worlds: . This is worth pausing on, because the count is where beginners most often go wrong: the worlds are ordered pairs, so and are two different worlds, not one.
A probability model assigns a number to every possible world.
B. The two axioms
Everything follows from two requirements on that assignment:
In words: no probability is negative or greater than one, and the probabilities of all the possible worlds add up to exactly one - something must happen.
That is the whole foundation. These axioms trace to Kolmogorov's Foundations of the Theory of Probability (1950), and every formula in the rest of this article is a consequence of them.
If the two dice are fair and do not interfere with each other, symmetry forces each of the 36 worlds to carry the same probability, and since they must sum to 1, each is .
C. Events: from worlds to propositions
We rarely care about a single world. We care about a proposition such as "the roll is even". An event is the set of worlds where the proposition holds, and its probability is the sum of their probabilities:
For one fair die and "even" :
Why "favourable over total" is a special case, not the definition. Counting outcomes and dividing works only when every world is equally likely. A loaded die still has a perfectly good probability model; it simply assigns unequal . The summation above is the real definition and always applies.
D. Deriving the complement rule
Here is the machinery working. The worlds where holds and the worlds where ("not ") holds are disjoint, and together they are all of . So splitting the second axiom's sum:
which is exactly
If a model gives an event probability , its non-occurrence has probability - and now you know that as a theorem, not an intuition.
E. The addition rule, and the trap it avoids
For two events that may overlap, adding probabilities double-counts the shared worlds. Removing the overlap exactly once gives
Worked example. Draw one card from a standard 52-card deck. Let = "it is a heart" and = "it is a King".
- - thirteen hearts.
- - four Kings.
- - exactly one King of Hearts.
Naïvely adding would have given , counting the King of Hearts twice. The overlap term is not a technicality; it is the difference between a right and a wrong answer.
Interactive: count the worlds, then count the overlap once
Every world is equally likely.
- P(A)
- 13/52 = 0.2500
- P(B)
- 4/52 = 0.0769
- P(A and B)
- 1/52 = 0.0192
- P(A or B)
- 16/52 = 0.3077
- Naive P(A) + P(B)
- 17/52 = 0.3269
- P(A) × P(B)
- 0.0192
Counting the worlds directly, A or B holds in 16 of 52. The addition rule gets there from the other three counts: 13 + 4 - 1 = 16, so P(A or B) = 4/13. The naive sum counts 17, and the extra 1 is exactly the highlighted overlap, counted once for A and again for B. These two events are also independent: P(A and B) equals P(A) × P(B), so knowing one tells you nothing about the other.
F. Conditional probability
Learning something changes what we should expect. The probability of given that holds is defined as
The intuition: restrict attention to the worlds where holds, then ask what fraction of that restricted world-set also has . The division by is what renormalises the restricted set back to total probability 1.
Rearranging gives the multiplication rule:
Worked example. In a population, 10% of members belong to a high-risk group. Within that group, an event occurs with probability 20%. The probability that a randomly chosen member is both high-risk and experiences the event is
that is, 2% of the whole population.
G. Independence, and why it must be earned
Two events are independent when knowing one tells you nothing about the other - formally, when . Substituting into the multiplication rule collapses it to the version most people remember:
For two genuinely unrelated events each of probability , the chance both occur is , or .
The most expensive mistake in applied probability. Multiplying probabilities is valid only under independence. When a common cause drives many outcomes at once - a shared failure mode, a market-wide shock, a single upstream dependency - the events are strongly dependent, and multiplying wildly understates the chance that many happen together. Independence is an assumption to be checked against the data-generating process, never assumed for free.
H. Verifying the die-and-card arithmetic
The numbers above are small enough to check by exhaustive enumeration, which is worth doing once so you trust the rules rather than the arithmetic:
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Running it prints 16 16 4/13: the union counted directly and the union computed
by the addition rule agree exactly, and the probability is the derived
above.
Key takeaways
- A probability model is a sample space of mutually exclusive, exhaustive possible worlds, each carrying a probability.
- The two axioms are that probabilities lie in and that they sum to 1 over . Everything else is derived.
- The probability of an event is the sum over the worlds where it holds; this is the definition, and "favourable over total" is only its equally-likely special case.
- "Or" needs the addition rule with the overlap subtracted; "and" needs the multiplication rule with a conditional probability.
- Independence is the special case where conditioning changes nothing. It must be verified, never assumed.
What's next
Conditional probability has a consequence far more useful than it first appears: it can be reversed, letting you turn "how the evidence behaves given a cause" into "how likely the cause is given the evidence". That reversal is Bayes' theorem, and it produces results that reliably surprise people. It is the subject of Bayes' Theorem and Belief Updating.
References & further reading
- Stuart Russell, Peter Norvig, Artificial Intelligence: A Modern Approach, Pearson (3rd edition), 2010· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.