The 95% Interval That Covers 81% of the Time
The textbook confidence interval for a proportion has exact coverage you can compute by summing over the n+1 possible samples, and at n = 30 with p = 0.10 it is 0.8085 rather than 0.95. Coverage does not improve monotonically with n, and in a rare-event setting it can fall to 0.0392. Two one-line alternatives fix it.
Prerequisites: What a Sample Can and Cannot Tell You
A 95% confidence interval promises one thing: over repeated samples, the interval contains the true parameter 95% of the time. For a proportion this is checkable exactly. There are only possible samples, so you can build the interval for each, ask whether it contains , and add up the binomial probabilities of the ones that do. No simulation, no approximation.
Do that for the standard interval, , at and :
One in five samples produces an interval that does not contain the truth. The interval is labelled 95%.
A. Not a small-sample caveat
The usual defence is that 30 is a small sample. Here is the same calculation across a range of settings, with two alternatives alongside:
| Wald | Wilson | Agresti-Coull | ||
|---|---|---|---|---|
| 30 | 0.10 | 0.8085 | 0.9742 | 0.9742 |
| 40 | 0.05 | 0.8681 | 0.9520 | 0.9861 |
| 100 | 0.05 | 0.8775 | 0.9659 | 0.9659 |
| 20 | 0.20 | 0.9208 | 0.9563 | 0.9563 |
| 50 | 0.30 | 0.9347 | 0.9567 | 0.9567 |
| 100 | 0.50 | 0.9431 | 0.9431 | 0.9431 |
A hundred observations of a 5% event still gives 0.8775. And the last row is worth noting on its own: at the most favourable proportion there is, with a hundred observations, the interval still covers 94.31% rather than 95%. It is never quite right, and in the corner of the parameter space where most real questions live - rare events, moderate samples - it is not close.
B. More data does not fix it monotonically
Coverage does not climb as grows. At :
| 25 | 30 | 35 | 40 | 45 | 50 | 55 | 60 | |
|---|---|---|---|---|---|---|---|---|
| coverage | 0.9187 | 0.8085 | 0.8713 | 0.9145 | 0.9356 | 0.8789 | 0.9153 | 0.9413 |
Going from 25 observations to 30 takes the coverage from 0.9187 down to 0.8085. Going from 45 to 50 takes it from 0.9356 down to 0.8789.
In the figure, pick 0.10 beside true p and drag sample size n to 30: Coverage at n reads 80.9%, the 0.8085 above.
Interactive: the coverage a 95% interval actually delivers
Exact, by enumeration. A binomial has only n + 1 outcomes.
- Coverage at n
- 95.1%
- Shortfall
- 0.0%
- Sizes that lose ground
- 29
- Samples with no successes
- 0.6%
At n = 23 the interval advertising 95% delivers 95.1%. Note the shape: coverage does not creep up toward the line, it jumps across it 29 times in this range alone, each tooth being one more attainable value of the estimate. Drag n from 23 to 24 and watch eight points disappear with a single extra observation. No rule of thumb about large enough n protects against that, because the thing being counted is discrete.
The oscillation is not noise, because there is no noise here: every entry is an exact sum. It happens because the set of achievable values is discrete, and as changes, one of the possible outcomes moves across the boundary of containing and takes its whole binomial probability with it. Any statement of the form "the interval is fine once is large enough" has to survive this table, and the usual rules of thumb (, and so on) do not.
C. The degenerate case
Take and . The coverage of the standard interval is
A 95% interval that contains the truth in 4% of samples. The mechanism is immediate: , and when the estimate is , the standard error is , and the interval is the single point , which does not contain . The formula reports perfect certainty in exactly the situation where it has seen nothing.
This is the case that turns up in production: the rare failure, the rare click, the rare adverse event.
D. What to use instead
Both alternatives in the table are one line and neither needs a new idea.
- Wilson. Invert the test rather than the estimate: keep the values of the data would not reject. The interval is no longer centred on and it never collapses to a point.
- Agresti-Coull. Add two successes and two failures, then use the ordinary formula on the adjusted counts. It is the Wald interval with the estimate pulled off the boundary, and in the table above it is the more conservative of the two.
The general lesson is the one the exact calculation makes available: coverage is computable, so compute it. For a proportion it is a finite sum. For anything else, simulating the sampling distribution and counting how often the interval contains the value you generated from takes a few lines, and it is the only way to know whether the number on the label is the number you get.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.