Understanding p-value
Begin with the assumption to be doubted. A hypothesis test starts by supposing the null hypothesis holds - no difference, no effect, nothing going on - and asks a single question of the data: if that were true, how often would a sample look at least this extreme? That probability is the p-value. Every part of its meaning follows from the fact that the null was assumed in order to compute it.
This is why the commonest reading is not merely imprecise but backwards. The p-value cannot be the probability that the null hypothesis is true, because the calculation already fixed the null as true; a quantity conditioned on an assumption cannot simultaneously be a probability about that assumption. Nor is it the probability that the result was a fluke, nor one minus the probability that the finding will replicate.
The distribution under the null makes the definition concrete. If the null holds and the test is valid, the p-value is uniformly distributed: every value between 0 and 1 is equally likely. Simulating 50,000 samples of size 25 from a true null and testing each gives a Kolmogorov-Smirnov statistic of 0.004299 against the uniform distribution, with 0.0499 of the p-values falling below 0.05. Nothing distinguishes a p-value of 0.03 from one of 0.53 except that someone declared the first region interesting.
What a small p-value does establish is that the data would be unusual if the null were true, which is genuine evidence and worth having. What it cannot do alone is tell you how much to believe the alternative, because that depends on how plausible the alternative was beforehand and on how large an effect the study could have detected. A p-value just under a threshold, from a small study, on a hypothesis that was unlikely to begin with, is weak evidence dressed as a discovery.
How to Calculate
p = P(T(X) ≥ t_obs | H₀), and under H₀ valid: p ~ Uniform(0,1)
where
- H₀
- the null hypothesis, assumed true throughout the calculation
- T(X)
- the test statistic computed from a hypothetical repeat of the study
- t_obs
- the value of that statistic actually observed in the data at hand
- ≥
- read as "at least as extreme as"; for a two-sided test the tail is taken on both sides
Example of p-value
Under a true null, 50,000 simulated two-sided t-tests on samples of 25 produce p-values whose Kolmogorov-Smirnov distance from Uniform(0,1) is 0.004299 (KS p-value 0.3129), and 0.0499 of them fall below 0.05. The 5% rate is the definition of the threshold rather than a property of the data.
Run 20 independent tests on hypotheses that are all null, and the probability that at least one comes out significant at 0.05 is 1 - 0.95^20 = 0.641514. Nothing has malfunctioned: each test is wrong 5% of the time as advertised, and 20 opportunities is enough for that to be likely at least once.
Suppose one hypothesis in ten is genuinely true and each test has power 0.8 at alpha 0.05. Per 100 hypotheses, 8 true effects are detected and 4.5 false ones are raised from the 90 nulls, so 0.36 of the significant results are false. Every test is behaving exactly as specified; the base rate did the damage.
Frequently Asked Questions
Does p < 0.05 mean the effect is real?
It means the data would be unusual if there were no effect. Whether the effect is real also depends on how plausible it was beforehand and how much power the study had. With 10% of tested hypotheses true and 80% power, more than a third of results significant at 0.05 are false positives.
Does a large p-value mean there is no effect?
No. It means the study did not find the effect, which is only informative in proportion to its power. A two-sided one-sample t-test at n = 25 detects an effect of half a standard deviation just 0.669708 of the time, and an effect of a fifth of a standard deviation only 0.160504 of the time.
Why 0.05?
Convention, and an explicitly arbitrary one. It sets the long-run rate of false positives among true nulls, so it is a statement about how often one is willing to be wrong. Fields that test many hypotheses at once routinely require far smaller thresholds for exactly that reason.
The Bottom Line
A p-value answers one narrow question well: how surprising would this data be if nothing were going on? It cannot answer the question people want, which is how much to believe the effect - that needs the plausibility of the hypothesis beforehand and the power of the study behind it.