Skip to content
Kudos AI

Confounding and the Reversal

A treatment that helps by the same margin in every subgroup while appearing to harm overall, the mechanism that produces it, and the standardisation that recovers the right answer.

IntermediateModule 125 min · 100 XP
Patients sorting by severity so the treated arm fills with the harder cases, within-group recovery bars staying five points apart while the pooled bars cross over, and the standardised comparison putting the sign back.

Start with a set of numbers that sounds impossible, and resist the urge to look for the arithmetic error. There isn't one.

The table

A treatment is given to 60 patients and withheld from 60 others.

treateduntreated
severe16 / 40 = 0.4000.4007 / 20 = 0.3500.350
mild14 / 20 = 0.7000.70026 / 40 = 0.6500.650
pooled30 / 60 = 0.5000.50033 / 60 = 0.5500.550

Among severe cases the treatment raises recovery by five points. Among mild cases it raises recovery by five points. Pooled, the treated do five points worse.

Every patient is in exactly one stratum. Both arms have 60 patients, so this is not a matter of unequal group sizes. The addition checks out.

What the pooled number is actually comparing

Look at who ended up in each arm.

  • of the 60 treated, 40 were severe - two thirds
  • of the 60 untreated, 20 were severe - one third

Doctors gave the drug to the patients who needed it, which is what any conscientious clinician would do. The consequence is that the treated arm is loaded with the harder cases.

Severe patients recover less often whatever is done to them. So the pooled comparison is not measuring the treatment at all - it is measuring a group that is mostly severe against a group that is mostly mild, and that difference in case mix is larger than the five points the treatment contributes.

Interactive: the same treatment, two verdicts

The four within-group rates never change.

treateduntreated
severe16 / 4040.0%7 / 2035.0%
mild14 / 2070.0%26 / 4065.0%
pooled30 / 6050.0%33 / 6055.0%
Pooled difference
-5.0
Standardised difference
+5.0
Pooled verdict
harms

The pooled figures say the treatment harms, by 5.0 points - and it helped in the severe group and helped in the mild group. Every patient is in exactly one of those two rows. Nothing in the top four cells has moved: the only thing you changed is who is in each arm. Two thirds of the treated are severe here and only one third of the untreated are, and severe patients recover less often whatever is done to them, so the pooled comparison is measuring the case mix and calling it a treatment effect.

Confounding, and the two arrows it needs

Severity did two things at once:

  1. it influenced who was treated
  2. it influences who recovers

A variable doing both is a confounder, and both arrows are required.

A variable affecting only the outcome - some genuine risk factor nobody consulted when prescribing - costs precision but does not bias the comparison, because it is spread evenly across the arms. A variable affecting only the assignment, such as a coin, is harmless for the same reason.

It is the common cause that does the damage, which is why the question worth asking about any observational comparison is how the groups came to be formed.

Standardising

Because severity was recorded, it can be adjusted for. The pooled figure averaged each arm's recovery over that arm's own case mix. Instead, average the within-stratum rates over one common mix:

treated: 0.5×0.400+0.5×0.700=0.550\text{treated: } 0.5 \times 0.400 + 0.5 \times 0.700 = 0.550 untreated: 0.5×0.350+0.5×0.650=0.500\text{untreated: } 0.5 \times 0.350 + 0.5 \times 0.650 = 0.500

A difference of exactly +0.05+0.05: the same five points visible in each stratum, and the opposite sign to the naive −0.05-0.05.

Why more data does not help

This is the point that separates a statistical problem from a causal one, and it is worth stating flatly.

Sampling noise shrinks as the sample grows. Confounding does not. It is a property of how the data came to exist, not of how much of it you have.

Run the same study on a million patients under the same prescribing habits, and the naive estimate converges - to −0.05-0.05. You get a narrow confidence interval around a number with the wrong sign, and every conventional check will look healthy.

Notice also what the data cannot tell you. The naive and the standardised estimates are both correct arithmetic on the same table. Nothing inside it announces which one answers the causal question. That judgement comes from knowing how treatment was assigned, which is information about the world rather than about the numbers.

What this sets up

Standardising worked here because severity was recorded. That is the fortunate case, and it rests on a claim that cannot be checked from inside the data: that severity was the only common cause.

Two questions follow, and they are the next two lessons. What if you cannot list the confounders - is there a way to defeat them without naming them? And once you start adjusting for variables, which ones should you adjust for, given that some adjustments make the answer worse?

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Unlock the full path

This first lesson is free. Enrol to take the mastery quiz, earn XP, and unlock every module, with more interactive, runnable examples throughout.