Skip to content
Kudos AI

Statistical Power

The probability that a test rejects the null hypothesis when a specified alternative is true. It is the chance of finding an effect that is genuinely there, and it is fixed by the design before any data are collected.

Also known as: Sensitivity of a test, 1 − β

Understanding Statistical Power

A test can be wrong in two ways, and they are not symmetric. It can reject a true null, which happens at the rate alpha that the analyst chooses, or it can fail to reject when an effect is genuinely present, which happens at a rate beta that the analyst mostly inherits from the design. Power is 1 - beta: the probability of detecting the effect.

Power is always power against something. It depends on the size of the effect being looked for, the variability of the measurements, the sample size and the chosen alpha, and quoting it without naming an effect size means nothing. The convention is to express the effect in standard deviations, so that "half a standard deviation" describes a difference that is meaningful across measurement scales.

The practical consequence is that a non-significant result is not a finding of no effect unless the study could plausibly have found one. A two-sided one-sample t-test looking for half a standard deviation has power 0.293176 at n = 10, 0.669708 at n = 25, 0.933898 at n = 50 and 0.998610 at n = 100. The study of 25 that reports nothing would have missed a real effect of that size a third of the time; at a fifth of a standard deviation the same design misses five times in six.

Underpowered work also damages the results that do reach significance. If only one hypothesis in ten is true and power is 0.8, then per 100 hypotheses the tests yield 8 true detections and 4.5 false alarms, so 0.36 of significant findings are wrong. Lowering power raises that fraction, because it shrinks the true discoveries in the denominator while leaving the false positives in the numerator untouched - which is why power is a question about the credibility of a literature, not only about one experiment.

How to Calculate

power = 1 − β = P(reject H₀ | H₁ true), with β = P(type II error)

where

α
the type I error rate: rejecting a true null, chosen by the analyst (commonly 0.05)
β
the type II error rate: failing to reject when the alternative holds
H₁
a specific alternative, including the effect size - power is undefined without one
δ√n
the noncentrality: effect size in standard deviations times the root sample size

Example of Statistical Power

Two-sided one-sample t-test at alpha 0.05, effect size half a standard deviation, computed from the noncentral t distribution rather than simulated: power 0.293176 at n = 10, 0.669708 at n = 25, 0.933898 at n = 50, 0.998610 at n = 100. Reaching the conventional 80% requires n = 34.

The same design looking for a fifth of a standard deviation has power 0.160504 at n = 25. Such a study fails to detect a real effect five times out of six, so its null result carries almost no information.

With 10% of tested hypotheses true, alpha 0.05 and power 0.8, the false discovery rate among significant results is (0.9 × 0.05) / (0.1 × 0.8 + 0.9 × 0.05) = 0.36. Running 20 independent tests on true nulls produces at least one significant result 0.641514 of the time; a Bonferroni threshold of 0.0025 brings that back to 0.048830.

Frequently Asked Questions

Can power be computed after the fact from the observed effect?

Not usefully. Power computed from the effect the study happened to observe is a restatement of the p-value rather than new information. Power is a design quantity, and the effect size it is computed against should be the smallest one worth detecting.

How much data does more power cost?

More than intuition suggests, because precision improves with the square root of the sample size. Going from power 0.669708 to 0.933898 for a half-standard-deviation effect means going from 25 observations to 50, and the returns keep shrinking after that.

Why not just lower alpha to be safer?

Because the two errors trade against each other. Lowering alpha reduces false positives and directly reduces power, raising the chance of missing real effects. Bonferroni at 0.0025 for 20 tests controls the family-wise error at 0.048830, and pays for it in every individual test.

The Bottom Line

Power is decided before any data exist, and it sets what both outcomes of the study will be worth: how much a null result licenses you to conclude, and what fraction of the significant results in a field are real. It is the cheapest thing to fix and the most expensive thing to discover afterwards.