Skip to content
Kudos AI
Lire en français
Causal Inference

The Control Variable That Invents a Relationship

Two independent causes and one common effect. Adjust for the effect and the causes acquire a correlation of exactly -1: a regression of A on B recovers a coefficient of +0.0030, and adding the common effect as a control turns it into -1.0000. Selecting a sample does the same thing invisibly, which is why "control for everything you measured" is not a defensible rule.

3 min readKudos AI

Prerequisites: The Treatment That Helps Everyone and Harms the Average

A three-node graph with two arrows meeting at one node, a cloud of points with no tilt at all, and the same cloud sliced at one value of the meeting node, every slice tilted the other way.

Here is the whole setup. Two causes, independent of each other, and one effect they both produce:

A⟶C⟵B.A \longrightarrow C \longleftarrow B.

CC is a collider: the arrows collide there. Nothing links AA to BB; whatever generated one had no knowledge of the other.

Now control for CC. Put it in the regression, or stratify on it, or - the case that catches people - collect a sample in which it is held fixed.

A. The binary version, exactly

Let AA and BB be independent fair coin flips, and let C=A or BC = A \text{ or } B. Three quarters of the population has C=1C = 1. Among them:

P(A=1∣B=0,C=1)=1.0000,P(A=1∣B=1,C=1)=0.5000.P(A = 1 \mid B = 0, C = 1) = 1.0000, \qquad P(A = 1 \mid B = 1, C = 1) = 0.5000.

The correlation between AA and BB is 00 in the population and −0.5000\mathbf{-0.5000} once you condition on C=1C = 1. The reasoning is plain enough to do in your head: if the alarm went off and it was not BB, it must have been AA. Knowing BB now tells you about AA because you have already been told that one of them fired.

B. The continuous version, and a regression

Let AA and BB be independent standard normals and C=A+BC = A + B. On 200,000 draws:

Regressioncoefficient on BB
AA on BB+0.0030+0.0030
AA on BB and CC−1.0000\mathbf{-1.0000}

Adding one control variable moves a coefficient from indistinguishable-from-zero to exactly minus one. It is not an artefact of the sample: the partial correlation is exact, because ρAC=ρBC=1/2\rho_{AC} = \rho_{BC} = 1/\sqrt{2} and

ρAB⋅C=0−121−12=−1.\rho_{AB \cdot C} = \frac{0 - \tfrac{1}{2}}{1 - \tfrac{1}{2}} = -1.

Restricting to the observations with CC near zero gives −0.9992-0.9992 on 5,606 points, which is the same thing done by selection rather than by adjustment.

C. Selection is conditioning, and it is invisible

The regression version at least leaves a trace: someone chose to add CC. The selection version leaves none, because the conditioning happened before the data arrived.

  • Admitted patients. Two independent conditions each raise the chance of admission. Among the admitted, they appear negatively associated, and one looks protective against the other.
  • Hired candidates. If interviews weigh experience and test scores, and the bar is a combination of the two, then among the hired the two are negatively correlated even if they are unrelated in the applicant pool. Every "our best engineers did not have the best scores" observation has this as its first candidate explanation.
  • Successful startups, published papers, surviving products. Any sample defined by a threshold on an outcome that several causes feed into.

The tell is always the same: the analysis population was chosen using something downstream of the variables being compared.

D. What follows for practice

  • "Control for everything you measured" is not a rule, it is a coin flip. Adjusting for a confounder removes bias; adjusting for a collider creates it. The two are indistinguishable in the data and distinguishable only in the causal structure you are willing to state.
  • Draw the graph before choosing the covariates. It need not be right in every detail to be useful; it needs to say which variables are upstream of the treatment and which are downstream of the outcome.
  • Never adjust for anything caused by the outcome or by the treatment. That single prohibition catches most of these cases, including the "post-treatment variable" family.
  • Say how the sample was selected, every time. If the selection rule involves an outcome, the analysis is conditioning on a collider whether or not anyone wrote a control variable down.

The figure puts both on one treatment-outcome graph, with a confounder Z and a collider C. Under Adjust for, choosing Z removes the bias, and adding C brings bias back.

Interactive: every control is a causal claim

The graph criterion and the arithmetic are computed separately. They always agree.

ZconfounderTtreatmentYoutcomeCcollider
Estimate
0.5400
True effect
0.3000
Error
0.2400
Back-door criterion
failed

Adjust for

The path T - Z - Y is still open, so non-causal association is still flowing through it and the estimate is off by 0.2400. Z raises the chance of treatment and raises the outcome by itself, so the treated group would have done better anyway. Adjusting for Z blocks the fork and the arithmetic lands exactly on the truth.

References & further reading

  • Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, An Introduction to Statistical Learning, with Applications in R, Springer (Springer Texts in Statistics 103), 2013source ↗

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

8 min readCausal Inference

The Treatment That Helps Everyone and Harms the Average

A treatment that raises recovery by exactly five points in every subgroup while appearing to lower it overall, why more data makes that conclusion more confident rather than more correct, what randomisation buys that adjustment cannot, and the case where controlling for a variable manufactures an association from nothing.

StatisticsMachine Learning
6 min readUnsupervised Learning

The Direction That Changes When You Change Units

Twelve people, two measurements, and three different first principal components: in millimetres the answer is almost pure height, in metres almost pure weight, and in centimetres an even blend - with the correlation fixed at 0.9500 throughout. What that says about what PCA maximises, why a proportion of variance explained of 99.999% can be a statement about metres rather than about people, and what standardising actually chooses.

Machine LearningStatistics
4 min readStatistical Inference

The 95% Interval That Covers 81% of the Time

The textbook confidence interval for a proportion has exact coverage you can compute by summing over the n+1 possible samples, and at n = 30 with p = 0.10 it is 0.8085 rather than 0.95. Coverage does not improve monotonically with n, and in a rare-event setting it can fall to 0.0392. Two one-line alternatives fix it.

Statistics
← Back to all articles