Skip to content
Kudos AI

Power and the Winner’s Curse

Why the sample size has to be fixed before the test starts, what an underpowered experiment reports when it does reach significance, and the reason halving the effect you care about quadruples the traffic you need.

IntermediateModule 125 min · 100 XP
A required sample size ballooning as the target effect shrinks, then a cloud of experiment estimates with the significance threshold sweeping across it, keeping only the exaggerated tail.

Two experiments are run on the same feature. One has 14,751 users per arm, finds a 1.0 point lift and ships. The other has 2,000 users per arm, finds a 2.4 point lift and ships enthusiastically.

The second team is more excited and more wrong, and the reason is arithmetic rather than luck.

Sizing the test

Suppose your conversion rate is 10% and you would care about a one point improvement. The per-arm sample needed to have an 80% chance of detecting it, at the usual 5% level, is

n≈(zα/22pˉ(1−pˉ)+zβp1(1−p1)+p2(1−p2))2(p2−p1)2n \approx \frac{\left(z_{\alpha/2}\sqrt{2\bar{p}(1-\bar{p})} + z_{\beta}\sqrt{p_1(1-p_1)+p_2(1-p_2)}\right)^2}{(p_2-p_1)^2}

which comes to 14,751 per arm. The shape that matters is in the denominator: the effect is squared.

effect you want to detectper-arm sample
2.0 points3,841
1.0 point14,751
0.5 points57,763

Halving the effect multiplies the traffic by 3.9. The standard error falls as 1/n1/\sqrt{n}, so resolving an effect half as large needs the error to halve, which takes four times the data.

That single fact reorganises how experiments get planned. The question is never "is there an effect" - there is almost always some effect. The question is what is the smallest effect worth detecting, because that number sets the cost of the entire test, and it has to be answered by whoever will act on the result.

What an undersized test reports

Now run the undersized version: 2,000 users per arm, against the same true one point lift. Its power is 17.8%, so it misses four times out of five. That part is expected and it is not the problem.

The problem is what happens on the runs that do reach significance. Across 20,000 simulated experiments, the average measured effect among the significant ones is 2.38 times the truth. And 0.8% of them are significant in the wrong direction.

Nothing is wrong with the estimator. Averaged over all 20,000 runs it is unbiased: the mean estimate is the true effect. What is biased is the average over the subset that passed a filter, and the filter is "the estimate was large enough to clear the threshold". At low power, clearing the threshold requires noise to have helped.

This is the winner's curse, and it is why an exciting result from a small test so reliably shrinks when someone repeats it at scale.

The figure below computes all of that in closed form rather than by simulation, so the exaggeration moves continuously as you drag the sample size. It opens on the undersized test: 2,000 per arm, 17.8% power, winners overstating the truth by 2.39 times, and 0.8% of them significant backwards. Walk the slider up to 14,751 and two things happen at once. The exaggeration collapses towards one, and the power read-out returns exactly 80.0%, which is the sizing formula and this calculation agreeing on the same threshold.

Interactive: what a small test reports when it wins

Exact, not simulated. The filter is "large enough to clear the bar".

3x1x50060,000
Power
17.8%
Exaggeration
2.39x
Significant, wrong sign
0.8%
Needed for 80%
14,751
base ratetrue lift

At 2,000 per arm this test finds the effect 17.8% of the time, and that is not the problem. The problem is the runs that do win: they report 2.39 times the true effect on average, and 0.8% of them are significant in the wrong direction entirely. Nothing is wrong with the estimator - over every run it is unbiased. What is biased is the average over the runs that passed a filter, and at this power clearing the bar requires noise to have helped. Calling the result preliminary does not repair it. Sizing it at 14,751 does.

Why "preliminary" is not a defence

The usual response is to call the small test preliminary and treat its number as indicative. That does not work, for a reason worth stating plainly.

An underpowered experiment is not a noisy version of a good one. It is a machine that, whenever it speaks at all, speaks with a badly inflated number and a perfectly healthy-looking confidence interval around it. The interval is correct in the sense that it would cover the truth 95% of the time across all runs - but you are not looking at all runs, you are looking at the ones that were interesting enough to report.

So the sequence is fixed:

  1. decide the minimum effect worth acting on, with whoever will act
  2. compute the sample size that detects it
  3. if the traffic does not exist, either wait, widen the population, reduce variance (the third lesson), or do not run the test

Not running a test is a legitimate outcome. Running one that can only mislead is not.

A note on the other direction

Power is symmetric in a way people forget. A test with 17.8% power that finds nothing has told you almost nothing: the absence of a significant result is entirely consistent with the effect you were hoping for. Reporting that as "no effect" is the same error in the opposite direction.

The honest report of a null result is the confidence interval. "The lift was 0.2 points, 95% interval −1.7 to 2.1" says clearly that the experiment could not distinguish a large gain from a large loss. "No significant difference" hides that entirely.

What this sets up

Sizing the test correctly protects you from one failure mode. The next lesson covers the one that survives any sample size: an experiment that is checked repeatedly, or read across enough metrics, will eventually produce a winner even when nothing whatever is happening.

References & further reading

  • Ron Kohavi, Diane Tang, Ya Xu, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press, 2020· Kudos AI reference library
  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.