Understanding Mutual Information
Mutual information asks how much the uncertainty about one variable falls once you see another. Because it is built from the joint distribution rather than from a fitted line, it answers that question for any form of dependence, and it is zero if and only if the two variables are independent. That last clause is what correlation cannot offer: a correlation of zero rules out a linear relationship and nothing else.
The cleanest demonstration is the exclusive or. Generate two independent bits and take their XOR as the target. Each input on its own has a correlation with the target of about +0.002 and a mutual information of 0.0000 bits, which is correct in both cases - neither input alone says anything. Taken as a pair they have a mutual information of 1.0000 bit, because the pair determines the target exactly. The information exists only in the combination, so any screening procedure that ranks features individually discards both of them.
Applied to a channel, the same quantity says what can be carried. A binary channel that flips each bit with probability 0.1 has a capacity of 1 - H(0.1) = 0.5310 bits per use, and a simulation of it measures 0.5329. At a flip probability of 0.5 the capacity is exactly zero, because the output distribution is then the same whatever was sent. At 0.9 the capacity returns to 0.5310: consistent lying is as informative as consistent truth.
The data processing inequality closes the subject off. Anything you compute from the channel output is a function of that output, so it cannot know more about the input than the output did. In the same simulation, erasing three received bits in ten drops the measured information from 0.5329 to 0.2763 bits. No decoder, no reconstruction and no depth of network recovers what the channel destroyed, which is the honest ceiling on what any model can extract from a given set of features.
How to Calculate
I(X;Y) = Σ p(x,y) log2( p(x,y) / (p(x) p(y)) ) = H(X) − H(X|Y); capacity of a BSC = 1 − H(flip)
where
- I(X;Y) ≥ 0
- zero exactly when X and Y are independent
- H(X|Y)
- the uncertainty left about X once Y has been seen
- H(flip)
- the binary entropy of the channel error rate
- X → Y → Z
- a chain in which I(X;Z) can never exceed I(X;Y)
Example of Mutual Information
Two independent bits and their XOR, over 200,000 draws: each input alone has correlation +0.0016 and mutual information 0.0000 bits with the target; the pair has 1.0000 bit.
A binary channel with a 10% error rate: capacity 1 − H(0.1) = 0.5310 bits per use, measured at 0.5329. At a 50% error rate the capacity is 0.0000.
Erasing three received bits in ten in the same simulation lowers the measured information from 0.5329 to 0.2763 bits, and no processing raises it.
Frequently Asked Questions
Should I use mutual information instead of correlation for feature selection?
It catches more, but it is harder to estimate: on continuous variables it depends on binning or a density estimate, and it is biased upward on small samples. Use it knowing that a high value on few data points can be an artefact of the estimator rather than a finding.
Does high mutual information mean one variable causes the other?
No. It is symmetric in its two arguments and says nothing about direction, so a common cause produces exactly the same reading as a direct effect. Causal questions need assumptions about how the data came to exist.
If a network is deep enough, can it recover information the sensor lost?
No, and the data processing inequality is the proof. Every layer is a function of the previous one, so the ceiling is set where the information was destroyed. Improving the features is the only thing that raises it.
The Bottom Line
Mutual information measures dependence of any shape, in bits, and it puts a hard ceiling on what can be extracted downstream. It is the quantity to reach for when correlation reports nothing and you suspect there is something there, and the theorem attached to it tells you when to stop looking.