The Experiment That Was Going to Win Anyway
A test with 2,000 users per arm reports effects 2.4 times too large. An A/A test checked ten times comes out significant 19% of the time. Twenty independent null metrics produce a winner 64% of the time and twelve null segments 46%. Four numbers, one cause, and the decisions that have to be made before the data arrives.
Prerequisites: Statistical Inference
Run an experiment where the two arms are identical. Same page, same code, users split at random, nothing to find.
- Read it once at the planned size: significant 5.0% of the time. Correct.
- Check it ten times and stop when it looks good: 19.3%.
- Read twenty independent metrics instead of one: 64.2%.
- Slice it into twelve segments: 46.4%.
None of those numbers involves an effect. They are what a well-behaved test does when the reading rule is not the one the guarantee was written for. And they are the reason most disagreements about experiment results are really disagreements about decisions that should have been made a month earlier.
First: the test that is too small to be honest
Suppose your conversion rate is 10% and a one point lift would matter. To have an 80% chance of detecting it, you need 14,751 users per arm.
| effect to detect | per-arm sample |
|---|---|
| 2.0 points | 3,841 |
| 1.0 point | 14,751 |
| 0.5 points | 57,763 |
Halving the effect multiplies the traffic by 3.9, because the standard error falls as and the required sample therefore scales as one over the square of the effect.
Now run it with 2,000 per arm instead. Power drops to 17.8%, so it misses four times out of five. That is expected, and it is not the interesting part.
The interesting part is what it reports on the runs it does catch. Across 20,000 simulated experiments, the average effect among the significant ones is 2.38 times the truth, and 0.8% of them are significant with the wrong sign.
The estimator is not biased. Averaged over all runs it lands on the true effect exactly. What is biased is the average over the runs that passed a filter, and the filter is "the estimate was big enough to clear the threshold" - which at low power requires noise to have helped. This is the winner's curse, and it is why a striking result from a small test so reliably shrinks when someone repeats it properly.
So "we'll treat it as preliminary" is not a defence. An underpowered experiment is not a noisy version of a good one; it is a machine that, whenever it speaks at all, speaks with a badly inflated number and a healthy-looking interval around it.
The figure below computes all of that in closed form rather than by simulation, so the exaggeration moves continuously as you drag the sample size. It opens on the undersized test: 2,000 per arm, 17.8% power, winners overstating the truth by 2.39 times, and 0.8% of them significant backwards. Walk the slider up to 14,751 and two things happen at once. The exaggeration collapses towards one, and the power read-out returns exactly 80.0%, which is the sizing formula and this calculation agreeing on the same threshold.
Interactive: what a small test reports when it wins
Exact, not simulated. The filter is "large enough to clear the bar".
- Power
- 17.8%
- Exaggeration
- 2.39x
- Significant, wrong sign
- 0.8%
- Needed for 80%
- 14,751
At 2,000 per arm this test finds the effect 17.8% of the time, and that is not the problem. The problem is the runs that do win: they report 2.39 times the true effect on average, and 0.8% of them are significant in the wrong direction entirely. Nothing is wrong with the estimator - over every run it is unbiased. What is biased is the average over the runs that passed a filter, and at this power clearing the bar requires noise to have helped. Calling the result preliminary does not repair it. Sizing it at 14,751 does.
Second: every extra look is another chance
The estimated difference wanders as data accumulates. The significance threshold is a fixed line. A wandering quantity crosses a fixed line eventually, and a rule that stops at the first crossing keeps the excursion while discarding the return that would have followed.
That is the whole mechanism, and it turns 5.0% into 19.3% at ten looks. Continuous monitoring pushes it towards certainty given enough time.
Two objections that come up and do not hold:
- it is not that early data is noisier and therefore wrong. Each look is a valid test on the data available. The invalid object is the rule combining them.
- it is not fixed by watching the trend instead of the p-value. Any rule that decides to stop based on the result so far has the same structure.
If you need to monitor - and for catching outages you do - use a sequential test or an alpha-spending schedule. Those spend the error budget deliberately. What breaks the guarantee is spending it by accident.
Third: the same arithmetic, sideways
for independent null metrics. Real metrics are usually positively correlated, which brings these rates down somewhat.
| null metrics | chance at least one is significant |
|---|---|
| 1 | 5.0% |
| 5 | 22.6% |
| 20 | 64.2% |
| 50 | 92.3% |
At twenty metrics it is more likely than not; at fifty it is nearly guaranteed. Segments behave the same way: twelve of them produce a winner in 46.4% of null experiments, and there is always a story afterwards about why mobile users or new users behave differently. The story is generated after the number, and it is generated just as fluently when the number is noise.
Note the part that surprises people: a larger sample does not fix segments. The rate depends on the number of chances, not their precision. A bigger experiment produces spurious segments just as often, with tighter intervals around them.
Bonferroni brings twenty metrics back to 4.9% by testing each at 0.25%, which is correct and costs a great deal of power. Holm is uniformly better at the same guarantee. Benjamini-Hochberg controls the share of false claims rather than the chance of any, which is the right target when screening candidates and the wrong one when a single false claim is expensive.
But the limitation that matters is not mathematical. A correction only counts the comparisons you declare. The analyst who checked twenty metrics, three segments and four windows before reporting the one that worked made hundreds of implicit comparisons that nothing recorded. The p-value is then a number about a procedure nobody followed.
Interactive: how many chances before something wins
Both arms identical. Nothing to find.
- At least one significant
- 64.2%
- After Bonferroni
- -
- Per-test level needed
- -
With 20 null comparisons at the 5% level, something comes out significant 64.2% of the time. Nothing is wrong with the data, the test or the analyst: this is 1 − 0.9520, and it was true before anyone looked. Two things this does not depend on are worth noticing. It does not depend on the sample size - a bigger experiment produces spurious segments just as often, with tighter intervals around them. And it counts the comparisons you made, not the ones you reported. Turn the correction on to see what fixing it costs.
The cheap fix: buy precision instead of chances
There is a legitimate way to make an experiment cheaper, and it is not any of the above.
Users differ enormously from each other, and most of that has nothing to do with your change. A heavy user was heavy last month too. If you have a pre-experiment measurement for each user, subtract off the part of the outcome it predicts:
applied identically to both arms. At a pre-period correlation of 0.7 this removes 50.3% of the variance - the theory says - so the same precision comes from half the traffic.
It cannot bias the result because is measured before assignment and cannot have been affected by the treatment. Adjusting for something measured during the experiment breaks exactly that, and can manufacture an effect that is not there.
Two more that cost nothing: trimming a heavy-tailed metric, decided in advance, and stratifying the assignment so the arms are balanced by construction rather than on average.
Neither of this lesson's two numbers needs the simulation that produced it, and the figure below computes both. The adjustment removes exactly rho squared of the variance, so 49% at a correlation of 0.7 and the 50.3% measured above is that figure with 1,500 runs of noise on it. And a clean slice into twelve segments turns up at least one significant result 45.96% of the time, which is what the 46.4% is estimating. Drag either control and watch the trade.
Interactive: the traffic a column you already had is worth
Both numbers exact. The lesson simulates them; they do not need it.
- Variance removed
- 49.0%
- Traffic still needed
- 51.0%
- Per arm, plain
- 14,751
- Per arm, adjusted
- 7,524
A pre-period measurement correlated 0.70 with the outcome removes 49.0% of the variance - exactly rho squared, because the adjusted outcome is the residual of Y on X. The same precision then comes from 51.0% of the users: 7,524 per arm against 14,751. That is the legitimate way to make an experiment cheaper, and it cannot bias the result because X was measured before assignment and the same coefficient is applied to both arms. Adjusting for something measured DURING the experiment breaks both of those.
And then: significance is not a decision
The experiment is sized, pre-registered and precise. It returns a significant 0.1% lift. Ship?
Significance answers whether the effect is distinguishable from zero, and with enough traffic almost anything is. The decision needs the interval against the cost:
- 0.02% to 0.18% is real and almost certainly not worth maintaining the feature
- 0.02% to 3.5%, a barely significant 1.8% lift, is a different situation: the experiment has not resolved the question
- a null result with an interval of −0.05% to +0.07% is genuinely informative - you have ruled out anything worth shipping
- a null result with an interval of −4% to +5% has told you nothing at all
That is why the minimum effect worth shipping is fixed at the same moment as the sample size, by the people who will act on it. Fixed in advance it makes the decision arithmetic. Left open it becomes an argument in which everyone reads the same interval in the direction they already preferred.
What the discipline actually is
Write it down before the data arrives:
- the minimum effect worth acting on, agreed with whoever will act
- the sample size that detects it, and the decision not to run if the traffic does not exist
- one primary metric, plus guardrails declared separately
- a fixed population, duration and stopping rule, with a sequential method if you need to stop early
- any variance reduction you intend to apply
- everything else labelled exploration, which produces hypotheses and never conclusions
None of the failures above is a statistical subtlety. Each is the consequence of deciding something after seeing the data that should have been decided before, and the entire discipline of experimentation is the practice of moving those decisions earlier.
References & further reading
- Ron Kohavi, Diane Tang, Ya Xu, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press, 2020· Kudos AI reference library
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.