The Two Features That Look Like Noise
A variable that determines another with a correlation of exactly 0.0000000000, and a pair of features whose every pairwise mutual information with the target is exactly zero while the two together determine it completely. Univariate screening discards both, and the second case is the one that matters: the features it removes are removed because they matter.
Prerequisites: Entropy and Information
Let be uniform on and let . Knowing tells you with certainty. The correlation between them is
exactly, by symmetry. Correlation measures one thing: how much of moves linearly with . Here none of it does, and all of it is determined.
Mutual information is not fooled:
Knowing removes all of 's uncertainty. This part of the story is familiar, and the usual conclusion is "use mutual information instead of correlation". The second construction is the one that should worry you, because mutual information does not save you from it either.
A. Three variables, no pair informative
Let and be independent fair bits and let , the exclusive or. Every pairwise mutual information:
Any two of the three are independent. Look at alone and looks like a coin flip. Look at alone and looks like a coin flip. Yet
which is all of . The pair determines the target completely and neither member of the pair carries a single bit about it.
Open the figure's second tab, the pair that hides together, for this XOR pair: each input alone carries 0 bits about the target, and the pair carries 1.
Interactive: what a channel carries, and what it never gets back
Exact, from the joint distribution. No sampling anywhere.
- Bits per use
- 0.5310
- After processing
- 0.3199
- Lost to processing
- 0.2111
- Composite flip
- 0.1800
A channel that flips with probability 0.10 is right 90% of the time and still carries only 0.5310 bits per use. Accuracy and information are not the same currency. Now send the output through a second channel at 0.10: the composite flip probability is 0.1800 and what survives is 0.3199 bits. It fell, and it always will. That is the data processing inequality, shown here with a cascade because the arithmetic is checkable, but it holds for any function of the received signal whatever.
B. What this does to feature selection
The standard first step on a wide dataset is univariate screening: score each feature against the target, keep the top ones, model those. Correlation, mutual information, a single-feature AUC, a chi-squared test - it does not matter which, they all agree here.
On the XOR data, every screening method scores and at exactly zero and drops both. A screen that kept the top 10% of a thousand features would drop them; a screen that kept the top 90% would drop them. They are not near the threshold, they are at the floor.
And the failure is not a rare corner. Any feature that acts through an interaction looks like this to some degree:
- a drug that helps one genotype and harms another, with no main effect
- a control that matters only above a temperature
- a difference between two measurements when neither measurement alone is informative
- any target that depends on a parity, a ratio, or a match between two fields
C. What to do instead
- Screen with a model that can see pairs. A shallow gradient-boosted tree on all features, scored by permutation importance, finds and immediately, because a depth-two tree can represent XOR and a single split cannot.
- Screen in blocks rather than singly, when the features have known structure: the two ends of a measurement, the before and after, the pair of fields that are compared downstream.
- Construct the interaction if you suspect one. A difference, a ratio or a product computed before screening turns a pairwise effect into a univariate one, and is checkable in a line.
- Treat "no feature was predictive" as a hypothesis, not a finding. The XOR table is what that conclusion looks like when it is wrong.
The general statement is worth keeping in the exact form the numbers give it. Independence of every pair does not imply independence of the set: the joint distribution of is not the product of its margins even though every two-way margin factorises. Every screening procedure that scores features one at a time is assuming the opposite.
References & further reading
- Thomas M. Cover, Joy A. Thomas, Elements of Information Theory, Wiley (2nd edition), 2006· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.