The Treatment That Helps Everyone and Harms the Average
A treatment that raises recovery by exactly five points in every subgroup while appearing to lower it overall, why more data makes that conclusion more confident rather than more correct, what randomisation buys that adjustment cannot, and the case where controlling for a variable manufactures an association from nothing.
Prerequisites: Probability and Statistical Foundations
Here is a set of numbers. All of them are exact, none is a trick, and together they say something that sounds impossible.
A treatment is given to 60 patients and withheld from 60 others. Among the severe cases it raises recovery from to . Among the mild cases it raises recovery from to . Five points, in both groups.
Pooled across everyone, the treated recover and the untreated recover . The treatment appears to harm.
Every patient belongs to exactly one of those two groups. The treatment helped in both. And it lowered the overall rate by exactly the same five points it raised each group by.
Where the numbers come from
| treated | untreated | |
|---|---|---|
| severe | 16 / 40 = | 7 / 20 = |
| mild | 14 / 20 = | 26 / 40 = |
| pooled | 30 / 60 = | 33 / 60 = |
Check the addition: out of , and out of . Both arms have 60 patients, so the pooled comparison is not being distorted by unequal sizes.
What differs is who is in each arm. Two thirds of the treated were severe; only one third of the untreated were. And severe patients recover less often whatever you do to them.
So the pooled figure is not comparing the treatment. It is comparing a group that is mostly severe against a group that is mostly mild, and the case mix overwhelms the five points the treatment contributes.
Interactive: the same treatment, two verdicts
The four within-group rates never change.
| treated | untreated | |
|---|---|---|
| severe | 16 / 4040.0% | 7 / 2035.0% |
| mild | 14 / 2070.0% | 26 / 4065.0% |
| pooled | 30 / 6050.0% | 33 / 6055.0% |
- Pooled difference
- -5.0
- Standardised difference
- +5.0
- Pooled verdict
- harms
The pooled figures say the treatment harms, by 5.0 points - and it helped in the severe group and helped in the mild group. Every patient is in exactly one of those two rows. Nothing in the top four cells has moved: the only thing you changed is who is in each arm. Two thirds of the treated are severe here and only one third of the untreated are, and severe patients recover less often whatever is done to them, so the pooled comparison is measuring the case mix and calling it a treatment effect.
The mechanism has a name and a shape
Severity influenced who got treated - doctors gave the drug to the patients who needed it - and it influences who recovers. A variable with both of those arrows is a confounder.
Both arrows are needed. A variable affecting only the outcome, such as a genuine risk factor nobody consulted when prescribing, costs you precision but does not bias the comparison. A variable affecting only the assignment, such as a coin, is equally harmless. The damage requires a common cause.
The repair, and what it gives
Since severity was recorded, it can be adjusted for. Compute the effect within each stratum, then average those over one common case mix instead of the mix each arm happened to have:
A difference of exactly - the same five points visible in each stratum, and the opposite sign to the naive .
Why this is not a sample-size problem
This is the part worth internalising, because it separates a statistical difficulty from a causal one.
Sampling noise shrinks as data accumulates. Confounding does not. It is a property of how the data came to exist, so collecting a million more patients under the same prescribing habits produces a tighter and tighter confidence interval around the wrong answer - and around the wrong sign.
Nothing computed from the table announces which of the two estimates is causal. The naive and adjusted numbers are both correct arithmetic on the same data. What distinguishes them is knowledge of how treatment was assigned, and that comes from outside the data.
What randomisation buys
Now take a similar world, not the one in the table: half the patients severe, untreated recovery for the severe and for the mild, and a true benefit fixed at exactly by construction. Work it out twice, changing only who decides on treatment - doctors who treat 80% of the severe and 20% of the mild, or a coin:
| assignment | naive estimate |
|---|---|
| doctors choose, by severity | |
| a coin chooses |
Those are the exact population values; a 1,500-trial simulation of the same world lands a few thousandths away, near and .
The treatment is identical in both. When doctors assign, severity flows into the treatment variable and the estimate reverses. When a coin assigns, no patient characteristic can influence the arm, and the naive comparison recovers the effect.
That is the whole argument for randomisation, and it is stronger than adjustment in a specific way: adjustment can only handle confounders you know about and have recorded, and the claim that your list is complete cannot be checked from inside the data. Randomisation breaks the arrow from every unmeasured cause at once, without needing to name any of them.
It is worth being precise about the promise. Randomisation balances the arms in expectation; one small trial can still be unlucky, which is why trials are sized deliberately and baseline tables are inspected. And it protects exactly one point in the chain: assignment. If patients drop out afterwards for reasons connected to their treatment, the groups that remain are confounded again, which is why analyses are specified as intention-to-treat.
The world is specified exactly enough that neither row needs simulating, and the figure below computes both: -0.08 when severity decides and +0.10 when a coin does, which are the exact values in the table above. The slider is the exercise. Bring the two assignment probabilities together and the first row becomes the second, because the bias is exactly the severity imbalance times the 0.3 gap in baseline recovery. Note what that never required: a fair coin. Any common probability does it.
Interactive: the same world, assigned twice
Closed form. The treatment is identical in every reading.
- Naive comparison
- -0.0800
- Adjusted for how ill
- +0.1000
- True effect
- +0.1000
- Imbalance in who is severe
- +0.6000
- Severe among treated
- 80%
- Severe among untreated
- 20%
Being severe now decides the assignment, so the treated arm is 80% severe against 20% in the other - an imbalance of +0.6000. The naive comparison reads -0.0800 against a true effect of +0.1000, and at the lesson’s setting it lands on the wrong side of zero entirely. The bias is exactly the imbalance times the 0.3 gap in baseline recovery, which is why pulling the slider to the middle removes it and nothing else has to change.
The adjustment that makes things worse
The lesson so far sounds like "adjust for what you can". That is where many analyses go wrong, because some adjustments create bias.
Take two causes that are genuinely independent - each occurring with probability , with no relationship whatsoever. Suppose a case is admitted to a study when either cause is present.
Among the admitted:
- the probability of the first cause is
- but given that the second cause is also present, it drops to
The joint probability among the admitted is , while the product of the marginals is . Independence is gone.
Nothing changed in the world. Conditioning on a common effect - a collider
- manufactured a dependence between causes that had none. Intuitively: once you know a case was admitted, learning that one cause was present explains the admission and makes the other less likely.
This is why "control for everything you measured" is not sound advice, and why adding variables to a regression is a causal claim rather than a safety measure.
The figure below checks a choice of controls from both directions at once. The graph side runs d-separation over the enumerated paths; the arithmetic side runs the adjustment formula over the joint distribution. Neither consults the other, and they never disagree. Add the collider and watch a criterion fail and an estimate move on the same click - then add the confounder as well, and watch the damage stay, because a wrong control is not cancelled by a right one.
Interactive: every control is a causal claim
The graph criterion and the arithmetic are computed separately. They always agree.
- Estimate
- 0.5400
- True effect
- 0.3000
- Error
- 0.2400
- Back-door criterion
- failed
Adjust for
The path T - Z - Y is still open, so non-causal association is still flowing through it and the estimate is off by 0.2400. Z raises the chance of treatment and raises the outcome by itself, so the treated group would have done better anyway. Adjusting for Z blocks the fork and the arithmetic lands exactly on the truth.
The rule, and its price
Drawing the assumptions as a graph makes the distinction readable. Three shapes cover most cases:
- fork (): a common cause. Adjust for - this is confounding.
- chain (): a mediator. Adjusting for removes part of the very effect you are measuring.
- collider (): a common effect. Adjusting for creates a spurious association.
The back-door criterion turns this into a rule: adjust for a set of variables that blocks every path into the treatment from behind and contains no descendant of the treatment, so that no mediator is adjusted away and no collider opened.
When the assumptions hold, the result is exact rather than approximate. In a world built so the true effect is exactly , the naive comparison gives - overstating by - and the back-door adjustment returns exactly. Both figures are computed in exact fractions with no sampling, so that gap is bias, not noise.
The price is honesty about what the arrows rest on. Correlation is symmetric and causation is not, so several different graphs can generate the same table of numbers. The graph cannot be read off the data; it is brought to the data from knowledge of how things work.
That is a feature. It forces the assumption into the open where it can be argued with, instead of leaving it implicit in a list of regression controls.
What to carry away
A comparison between the treated and the untreated answers the causal question only if the groups were otherwise alike. When they were not, the estimate is not merely imprecise - it can point the wrong way, and more data will make it more confident rather than more correct.
So the most useful question about any observational result is not which test was run or how large the sample was. It is: how were these groups formed?
The training path Causal Inference works through all three arguments in detail, with every computation runnable and editable in the browser.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.