Skip to content
Kudos AI

Confounding

A variable that influences both the treatment and the outcome, so that a comparison between the treated and the untreated measures the difference between the groups as well as the effect of the treatment.

Also known as: Confounder, Common cause

Understanding Confounding

The comparison everyone reaches for is between those who received a treatment and those who did not. It answers the causal question only if the two groups were otherwise alike, and they usually are not, because something decided who received it. When that something also affects the outcome, the comparison is contaminated by it.

The mechanism is easiest to see when the numbers reverse. Suppose a treatment raises recovery from 0.350 to 0.400 among severe cases, and from 0.650 to 0.700 among mild ones - exactly five points in each. Pooled, the treated recover 0.500 against 0.550 for the untreated, and the treatment appears to harm. Nothing is miscalculated: two thirds of the treated were severe against one third of the untreated, so the pooled figure is comparing case mixes.

The repair for a confounder you have measured is standardisation: compute the effect within each stratum, then average those over one common distribution of the confounder rather than over the distribution each arm happened to have. Here that gives 0.550 against 0.500, a difference of exactly +0.05 - the same five points visible in each stratum, and the opposite sign to the naive -0.05.

Two consequences follow. First, confounding is not a small-sample problem: it is a property of how the data came to exist, so more of the same data tightens an interval around the wrong number. Second, no statistic computed from the table announces which estimate is causal. That judgement comes from outside, from knowing how treatment was assigned, which is exactly why randomisation is valuable and why observational studies argue about assumptions rather than about arithmetic.

How to Calculate

P(Y | do(T=t)) = Σ_z P(Y | T=t, Z=z) P(Z=z) ≠ P(Y | T=t)

where

T
the treatment, whose effect is wanted
Y
the outcome
Z
the confounder: a common cause of T and Y
do(T=t)
setting the treatment by intervention rather than observing it

Example of Confounding

Severe cases: 16 of 40 treated recover (0.400) against 7 of 20 untreated (0.350). Mild cases: 14 of 20 treated recover (0.700) against 26 of 40 untreated (0.650). Treatment helps by exactly 0.05 in each.

Pooled: 30 of 60 treated (0.500) against 33 of 60 untreated (0.550). The naive difference is -0.05, the standardised difference is +0.05, and the two have opposite signs.

In a world whose true benefit is +0.1000 by construction, letting doctors assign treatment by severity yields a naive estimate of -0.0800, while assigning by coin yields +0.1000.

Frequently Asked Questions

Does adjusting for more variables always reduce bias?

No. Adjusting for a common effect of two variables - a collider - creates an association that was not there, and adjusting for a variable on the causal path from treatment to outcome removes part of the effect you are trying to measure. Which variables to adjust for is a question about the causal structure, not about the data.

How do I know whether my estimate is confounded?

Not from the data. Two worlds, one confounded and one not, can produce identical tables. What distinguishes them is how treatment was assigned, so the useful question is always how the groups came to be formed.

Is confounding the same as correlation without causation?

It is one common reason for it, and an important one, but not the only one. Reverse causation and selection into the sample produce the same symptom, and they are repaired differently.

The Bottom Line

Confounding turns a comparison into a comparison of two different populations. It is bias rather than noise, so no amount of extra data removes it, and it can be strong enough to reverse a sign - which is why "how were these groups formed?" is a more useful question than any test computed from the numbers.