What a Sample Can and Cannot Tell You
Estimators as random variables with distributions of their own, the case where the unbiased estimator is the worse one, what a confidence interval actually promises and the standard interval that delivers 87% where it advertises 95%, and what a p-value is a probability of - every figure computed exactly or by fixed-seed simulation.
Prerequisites: Probability and Statistical Foundations
Almost everything quantitative you will ever read rests on one move: someone measured a few things and said something about the many things they did not measure. A trial of 200 patients becomes a claim about a treatment. A survey of 1,000 voters becomes a number on the news.
That move is inference, and it is not magic. It has rules, it has a cost, and it has three places where the conclusion routinely gets stated more strongly than the arithmetic allows. This article walks through those three, and every figure below was computed before it was written - exactly where an exact answer exists, and by simulation with a stated seed where it does not.
An estimate is itself a random thing
Start with the idea that makes the rest possible. When you compute an average from a sample, that average is not a fact about the world; it is a fact about your sample. Draw a different sample and you get a different number. So the estimate has a distribution of its own, and understanding inference means thinking about that distribution rather than about the single number you happened to get.
Two things describe it. Bias is whether the estimator is centred on the truth
- whether, averaged over all the samples you might have drawn, it lands on the right answer. Variance is how much it bounces around from sample to sample.
The classic illustration is the formula for variance. If you measure the spread of your data around the sample mean and divide by , you get an answer that is systematically too small - by exactly the factor , which at is nine tenths. The reason is worth pausing on: the sample mean is, by construction, the point that sits closest to your particular data. Measuring spread from it therefore understates the spread around the true mean, which is somewhere else. Dividing by restores exactly what was spent locating the centre.
Simulating 400,000 samples of size 10 from a normal population with variance 4 confirms it: the version averages against a prediction of , and the version averages .
Where unbiased turns out to be worse
Here is the part that is usually skipped, and it is the more useful half.
Being right on average is not the same as being close. The natural measure of "close" is mean squared error, which decomposes exactly into bias squared plus variance. On those same 400,000 samples:
- the biased estimator has mean squared error
- the unbiased estimator has
The unbiased estimator is the worse of the two, and not by a little. These are not simulation artefacts - for a normal population the exact values are and , and the simulation lands within a quarter of a percent of both. Dividing by the larger number shrinks every estimate slightly toward zero, and the variance that buys is worth more than the bias it costs.
This is the bias-variance trade-off in its purest form, and it is the reason methods that deliberately introduce bias - ridge regression, shrinkage, almost all regularisation - are not compromises but improvements.
One more consequence of the distribution: its width shrinks as . With a population standard deviation of 2, the standard error of the mean is at and at . Halving your uncertainty costs four times the data. That single fact governs the budget of every study ever designed.
All four of those numbers have closed forms, so the figure below computes them rather than simulating: 3.6 and 4 for the two recipes on average, 3.04 and 3.5556 for their mean squared errors. Dragging the sample size settles something this section states at one size only - for normal data the biased recipe is closer at EVERY n, because (2n-1)(n-1) is less than 2n squared for every n, and the advantage shrinks away as the sample grows.
Interactive: the recipe that is wrong and closer anyway
Closed form throughout. The lesson simulates these; they do not need it.
- Divide by n, on average
- 3.6000
- Divide by n - 1
- 4.0000
- Its mean squared error
- 3.0400
- And the other one’s
- 3.5556
Dividing by n recovers 3.6000 of a true variance of 4 on average - exactly 10 - 1 over 10 of it, because the sample mean is the one point that makes those squared distances as small as they can be. Dividing by n - 1 gives the degree of freedom back and lands on 4 exactly. And yet the biased recipe is CLOSER: its mean squared error is 3.0400 against 3.5556, split into a bias of -0.4000 and a spread of 2.8800. Drag n and the ordering never changes, at any sample size, because (2n-1)(n-1) is less than 2n^2 for every n. Both mean-squared-error formulas assume a normal population, which is where the fourth moment comes from.
What a confidence interval actually promises
A single number hides its own precision, so we report a range instead. The promise attached to a 95% confidence interval is precise, and it is not the one most people state.
The promise is about the procedure: build intervals this way, sample after sample, and 95% of them will contain the true value. It is not a statement about the interval on your screen. Once the data are collected and the numbers fixed, your interval either contains the truth or it does not - there is nothing left to be uncertain about. Saying "there is a 95% chance the true value is in here" treats a fixed unknown as if it were random.
That distinction sounds pedantic until you see what happens when the procedure is slightly wrong.
Two ways the standard recipe misses
Using the wrong quantile. When you do not know the population spread and must estimate it from the same small sample, there are two sources of uncertainty, and the familiar from the normal distribution accounts for only one. At , an interval built with it covers the truth of the time, not . That figure is exact, not simulated, for a normal population; on strongly skewed data both quantiles are approximations. The distribution exists precisely to fix this, and its quantile at 9 degrees of freedom is - noticeably wider. By 29 degrees of freedom it has already fallen to .
Estimating the width from the same slip. The textbook interval for a proportion is worse, and it fails in an instructive way. Because a binomial has finitely many outcomes, its coverage can be computed exactly by enumerating all of them. At :
| true | standard interval | Wilson interval |
|---|---|---|
| 0.05 | 0.919867 | 0.962224 |
| 0.10 | 0.878917 | 0.970308 |
| 0.20 | 0.937531 | 0.950701 |
| 0.50 | 0.935091 | 0.935091 |
At the advertised 95% is actually 87.9%. The mechanism is that the same unlucky sample both moves the centre and mis-sizes the width, so the two errors reinforce instead of cancelling.
At it collapses completely, covering . The reason is stark: of samples of 50 contain no successes at all, and when the estimate is zero the estimated standard error is also zero, so the interval is the single point - which misses.
The part that should end the rules of thumb
You might hope this is what "check that " protects against. It is not, because coverage does not improve smoothly as data accumulate. Still at , the standard interval covers:
- at
- at
One more observation costs nearly eight points of coverage. Here is 4.6 and 4.8, so the rule would have flagged both sizes, but it passes (), where coverage still falls six points from , from to . Coverage falls from one size to the next at 29 of the 180 steps between 20 and 200: most falls are one to three points, and three exceed five. A binomial with outcomes has a completely different set of achievable intervals when changes, so coverage jumps as whole outcomes cross the boundary. There is no threshold on that makes the standard interval safe, which is the honest argument for using a better one.
The figure below runs that enumeration for every sample size from 20 to 200 at once. Nothing is simulated: each point is a finite sum over the n + 1 outcomes. Start at n = 23, where the textbook interval covers 0.951 and looks blameless, then drag one step to 24 and watch eight points vanish. Switch to p = 0.01 to see the failure at its starkest, and to the Wilson interval to see what taking the width from p rather than from the estimate buys.
Interactive: the coverage a 95% interval actually delivers
Exact, by enumeration. A binomial has only n + 1 outcomes.
- Coverage at n
- 95.1%
- Shortfall
- 0.0%
- Sizes that lose ground
- 29
- Samples with no successes
- 0.6%
At n = 23 the interval advertising 95% delivers 95.1%. Note the shape: coverage does not creep up toward the line, it jumps across it 29 times in this range alone, each tooth being one more attainable value of the estimate. Drag n from 23 to 24 and watch eight points disappear with a single extra observation. No rule of thumb about large enough n protects against that, because the thing being counted is discrete.
What a p-value is a probability of
The third idea is the most misread, and the misreading follows from forgetting how it is computed.
A hypothesis test begins by assuming the thing you doubt. Suppose there is no effect. Then ask: if that were true, how often would data look at least as extreme as mine? That probability is the p-value.
Everything follows from the assumption being made first. The p-value cannot be the probability that the null hypothesis is true, because the calculation already fixed the null as true. A quantity conditioned on an assumption cannot also be a probability about that assumption.
Under a true null, the p-value is uniform on : every value is equally likely. Simulating 50,000 tests on samples of 25 from a true null gives a Kolmogorov-Smirnov distance of from the uniform distribution, with of the p-values below . A p-value of is no rarer than one of ; it merely falls in a region someone declared interesting.
Power decides what a null result is worth
The companion quantity is power: the probability of finding an effect that is genuinely there. It is fixed by the design, before any data exist, and it is always power against a specific effect size. For a two-sided one-sample (or paired) t-test on observations, looking for half a standard deviation, computed from the noncentral rather than simulated:
| power | |
|---|---|
| 10 | 0.293176 |
| 25 | 0.669708 |
| 50 | 0.933898 |
| 100 | 0.998610 |
Reaching the conventional 80% needs . So a study of 25 reporting "no significant effect" would have missed a real effect of that size a third of the time. Looking for a fifth of a standard deviation, the same design has power - it misses five times in six. Absence of evidence is evidence of absence only in proportion to power.
Why most significant findings can still be wrong
Now put the two together, with no test misbehaving anywhere.
Run 20 independent tests on hypotheses that are all null. The probability that at least one comes out significant at is . Nothing broke: each test is wrong 5% of the time exactly as advertised, and 20 chances make that likely.
Worse, suppose only one hypothesis in ten is genuinely true, with power and . Per 100 hypotheses tested, the tests find 8 real effects and raise 4.5 false alarms from the 90 nulls. So
More than a third of the significant results are false, while every individual test performs exactly to specification. What did the damage is the base rate - how many of the hypotheses were ever true. Bonferroni's threshold of pulls the family-wise error back to , and pays for it with power on every test.
The figure below lays a hundred hypotheses out as squares and runs the design you choose over all of them. Power is not a slider: it is computed from the sample size and the effect sought, through the noncentral t, so the curve beside the grid is the same table as above with the gaps filled in. Shrink the base rate and the amber squares - nulls the test called significant - start to outnumber the green, with nothing misapplied anywhere. Then tighten the threshold and watch the protection get paid for out of the power.
Interactive: a hundred hypotheses, every test behaving perfectly
Nothing here is misapplied. Count the amber squares anyway.
One square per hypothesis tested
- Power
- 80.8%
- False discovery rate
- 35.8%
- Significant and right
- 64.2%
- Real effects missed
- 1.9
Power against this effect, by sample size
Of the 12.6 findings this batch calls significant, 4.5 are false - 35.8% of them. No test misbehaved: each null was rejected at exactly its stated rate, and each real effect was found at exactly the power the design bought. What did the damage is the base rate, a quantity that appears nowhere in any of the tests. Raising power helps, but only through the numerator - it cannot touch the false alarms, which are set by the significance level and the number of nulls alone.
The through-line
The three ideas share a shape. An estimate, an interval and a p-value are each answers to a narrow, well-posed question, and each is routinely read as an answer to a broader one that people actually care about.
The estimate says where the data point, not how close it is. The interval describes a procedure's long-run behaviour, not your interval's chance of being right. The p-value says how surprising the data would be if nothing were going on, not how much to believe that something is.
Reading each as what it is costs nothing, and it is the difference between using statistics and being used by them.
The training path Statistical Inference works through all three in detail, with the computations available to run and change in the browser.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.