Skip to content
Kudos AI

Causal Graph

A drawing of assumed cause-and-effect relationships as arrows between variables, used to decide which variables must be adjusted for and which must not - a question the data alone cannot answer.

Also known as: Directed acyclic graph, DAG, Structural causal model

Understanding Causal Graph

A graph makes causal assumptions explicit. Each arrow asserts that one variable influences another, and the absence of an arrow is a claim too - often the stronger one. Because correlation is symmetric and causation is not, no table of associations orients the arrows on its own, and several distinct graphs can generate exactly the same numbers. Writing the graph down is what turns an assumption into something that can be argued with.

Three patterns cover most of the reasoning. In a chain, T causes M which causes Y, and controlling for M removes part of the effect you were trying to measure. In a fork, Z causes both T and Y, and controlling for Z is what removes confounding. In a collider, T and Y both cause C, and here the advice reverses: controlling for C manufactures an association between causes that were independent.

The collider case deserves its own arithmetic, because it contradicts the habit of adjusting for everything recorded. Take two independent causes, each present with probability one half, and admit a case when either is present. Among the admitted, the probability of the first cause is 2/3; among those admitted who also have the second, it drops to 1/2. Independence is gone: the joint probability among the admitted is 1/3 while the product of the marginals is 4/9. Knowing one cause explains the admission and makes the other less likely.

The back-door criterion reads the rule off the drawing: adjust for a set of variables that blocks every path into the treatment from behind and contains no descendant of the treatment, so that no mediator is adjusted away and no collider opened. When the assumptions hold, the adjustment is exact rather than approximate - in a world constructed so the true effect is 0.200000, the naive comparison gives 0.375325 and the adjusted estimate returns 0.200000 exactly. When an unmeasured common cause exists, the same formula returns a confident and wrong number, which is why the graph is the argument and the formula merely follows it.

How to Calculate

chain T → M → Y, fork T ← Z → Y, collider T → C ← Y

where

fork
a common cause; adjusting for Z is what removes confounding
chain
a mediator; adjusting for M removes part of the effect itself
collider
a common effect; adjusting for C creates a spurious association
back-door
the rule: block every non-causal path into T, adjusting for nothing downstream of T

Example of Causal Graph

Two independent causes of admission, each with probability 1/2: among the admitted, P(first) = 2/3, but P(first given second) = 1/2. The joint among the admitted is 1/3 against a product of marginals of 4/9, so conditioning destroyed the independence it did not find.

A confounded world built so the true effect is exactly 0.200000: the naive difference in outcomes is 0.375325, overstating it by 0.175325, and the back-door adjusted estimate is 0.200000 exactly. Both are computed in exact fractions, so the gap is bias rather than sampling noise.

The Simpson reversal is a fork: severity causes both the treatment decision and recovery, so the pooled comparison carries the wrong sign until it is standardised over severity.

Frequently Asked Questions

Can the graph be learned from data?

Partially, and never fully. Algorithms can narrow the possibilities to a class of graphs sharing the same conditional independences, but distinguishing within that class needs interventions or assumptions from outside the data.

What if I am unsure about an arrow?

Draw both versions and check whether the conclusion changes. If it does, the disagreement is the finding and should be reported as such; if it does not, the uncertainty was irrelevant. That is more informative than choosing one graph silently.

Is a regression with many controls equivalent to a graph?

No. A regression adjusts for whatever is in it, which the graph would tell you may include colliders and mediators. The controls encode a causal claim whether or not anyone writes it down, and the graph is that claim made legible.

The Bottom Line

A causal graph is the place where assumptions are stated rather than smuggled. It says which adjustments help, which do nothing, and which actively manufacture bias - and because different graphs can fit the same data, drawing it is the step that makes a causal claim arguable.