The Control Variable That Invents a Relationship
Two independent causes and one common effect. Adjust for the effect and the causes acquire a correlation of exactly -1: a regression of A on B recovers a coefficient of +0.0030, and adding the common effect as a control turns it into -1.0000. Selecting a sample does the same thing invisibly, which is why "control for everything you measured" is not a defensible rule.
Prerequisites: The Treatment That Helps Everyone and Harms the Average
Here is the whole setup. Two causes, independent of each other, and one effect they both produce:
is a collider: the arrows collide there. Nothing links to ; whatever generated one had no knowledge of the other.
Now control for . Put it in the regression, or stratify on it, or - the case that catches people - collect a sample in which it is held fixed.
A. The binary version, exactly
Let and be independent fair coin flips, and let . Three quarters of the population has . Among them:
The correlation between and is in the population and once you condition on . The reasoning is plain enough to do in your head: if the alarm went off and it was not , it must have been . Knowing now tells you about because you have already been told that one of them fired.
B. The continuous version, and a regression
Let and be independent standard normals and . On 200,000 draws:
| Regression | coefficient on |
|---|---|
| on | |
| on and |
Adding one control variable moves a coefficient from indistinguishable-from-zero to exactly minus one. It is not an artefact of the sample: the partial correlation is exact, because and
Restricting to the observations with near zero gives on 5,606 points, which is the same thing done by selection rather than by adjustment.
C. Selection is conditioning, and it is invisible
The regression version at least leaves a trace: someone chose to add . The selection version leaves none, because the conditioning happened before the data arrived.
- Admitted patients. Two independent conditions each raise the chance of admission. Among the admitted, they appear negatively associated, and one looks protective against the other.
- Hired candidates. If interviews weigh experience and test scores, and the bar is a combination of the two, then among the hired the two are negatively correlated even if they are unrelated in the applicant pool. Every "our best engineers did not have the best scores" observation has this as its first candidate explanation.
- Successful startups, published papers, surviving products. Any sample defined by a threshold on an outcome that several causes feed into.
The tell is always the same: the analysis population was chosen using something downstream of the variables being compared.
D. What follows for practice
- "Control for everything you measured" is not a rule, it is a coin flip. Adjusting for a confounder removes bias; adjusting for a collider creates it. The two are indistinguishable in the data and distinguishable only in the causal structure you are willing to state.
- Draw the graph before choosing the covariates. It need not be right in every detail to be useful; it needs to say which variables are upstream of the treatment and which are downstream of the outcome.
- Never adjust for anything caused by the outcome or by the treatment. That single prohibition catches most of these cases, including the "post-treatment variable" family.
- Say how the sample was selected, every time. If the selection rule involves an outcome, the analysis is conditioning on a collider whether or not anyone wrote a control variable down.
The figure puts both on one treatment-outcome graph, with a confounder Z and a collider C. Under Adjust for, choosing Z removes the bias, and adding C brings bias back.
Interactive: every control is a causal claim
The graph criterion and the arithmetic are computed separately. They always agree.
- Estimate
- 0.5400
- True effect
- 0.3000
- Error
- 0.2400
- Back-door criterion
- failed
Adjust for
The path T - Z - Y is still open, so non-causal association is still flowing through it and the estimate is off by 0.2400. Z raises the chance of treatment and raises the outcome by itself, so the treated group would have done better anyway. Adjusting for Z blocks the fork and the arithmetic lands exactly on the truth.
References & further reading
- Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.