Skip to content
Kudos AI
Lire en français
Probability Foundations

The Two Features That Look Like Noise

A variable that determines another with a correlation of exactly 0.0000000000, and a pair of features whose every pairwise mutual information with the target is exactly zero while the two together determine it completely. Univariate screening discards both, and the second case is the one that matters: the features it removes are removed because they matter.

3 min readKudos AI

Prerequisites: Entropy and Information

Two columns of values with a scatter that shows no tilt at all, and a joint table beside it whose cells are plainly not the product of its margins.

Let XX be uniform on {−2,−1,1,2}\{-2, -1, 1, 2\} and let Y=∣X∣Y = |X|. Knowing XX tells you YY with certainty. The correlation between them is

ρ(X,Y)=0.0000000000,\rho(X, Y) = 0.0000000000,

exactly, by symmetry. Correlation measures one thing: how much of YY moves linearly with XX. Here none of it does, and all of it is determined.

Mutual information is not fooled:

H(X)=2 bits,H(Y)=1 bit,I(X;Y)=1 bit.H(X) = 2 \text{ bits}, \qquad H(Y) = 1 \text{ bit}, \qquad I(X;Y) = 1 \text{ bit}.

Knowing XX removes all of YY's uncertainty. This part of the story is familiar, and the usual conclusion is "use mutual information instead of correlation". The second construction is the one that should worry you, because mutual information does not save you from it either.

A. Three variables, no pair informative

Let AA and BB be independent fair bits and let C=A⊕BC = A \oplus B, the exclusive or. Every pairwise mutual information:

I(A;B)=I(A;C)=I(B;C)=0.0000000000 bits.I(A;B) = I(A;C) = I(B;C) = 0.0000000000 \text{ bits}.

Any two of the three are independent. Look at AA alone and CC looks like a coin flip. Look at BB alone and CC looks like a coin flip. Yet

I(A,B ; C)=1 bit,I(A, B \,;\, C) = 1 \text{ bit},

which is all of CC. The pair determines the target completely and neither member of the pair carries a single bit about it.

Open the figure's second tab, the pair that hides together, for this XOR pair: each input alone carries 0 bits about the target, and the pair carries 1.

Interactive: what a channel carries, and what it never gets back

Exact, from the joint distribution. No sampling anywhere.

1 bit0.5
Bits per use
0.5310
After processing
0.3199
Lost to processing
0.2111
Composite flip
0.1800

A channel that flips with probability 0.10 is right 90% of the time and still carries only 0.5310 bits per use. Accuracy and information are not the same currency. Now send the output through a second channel at 0.10: the composite flip probability is 0.1800 and what survives is 0.3199 bits. It fell, and it always will. That is the data processing inequality, shown here with a cascade because the arithmetic is checkable, but it holds for any function of the received signal whatever.

B. What this does to feature selection

The standard first step on a wide dataset is univariate screening: score each feature against the target, keep the top ones, model those. Correlation, mutual information, a single-feature AUC, a chi-squared test - it does not matter which, they all agree here.

On the XOR data, every screening method scores AA and BB at exactly zero and drops both. A screen that kept the top 10% of a thousand features would drop them; a screen that kept the top 90% would drop them. They are not near the threshold, they are at the floor.

And the failure is not a rare corner. Any feature that acts through an interaction looks like this to some degree:

  • a drug that helps one genotype and harms another, with no main effect
  • a control that matters only above a temperature
  • a difference between two measurements when neither measurement alone is informative
  • any target that depends on a parity, a ratio, or a match between two fields

C. What to do instead

  • Screen with a model that can see pairs. A shallow gradient-boosted tree on all features, scored by permutation importance, finds AA and BB immediately, because a depth-two tree can represent XOR and a single split cannot.
  • Screen in blocks rather than singly, when the features have known structure: the two ends of a measurement, the before and after, the pair of fields that are compared downstream.
  • Construct the interaction if you suspect one. A difference, a ratio or a product computed before screening turns a pairwise effect into a univariate one, and is checkable in a line.
  • Treat "no feature was predictive" as a hypothesis, not a finding. The XOR table is what that conclusion looks like when it is wrong.

The general statement is worth keeping in the exact form the numbers give it. Independence of every pair does not imply independence of the set: the joint distribution of (A,B,C)(A, B, C) is not the product of its margins even though every two-way margin factorises. Every screening procedure that scores features one at a time is assuming the opposite.

References & further reading

  • Thomas M. Cover, Joy A. Thomas, Elements of Information Theory, Wiley (2nd edition), 2006· Kudos AI reference library

Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.

Related reading

7 min readInformation Theory

The Bound That Is Actually Reached

Entropy is not a summary of a distribution but a floor that the best code meets to the last decimal, the surcharge for using the wrong distribution is exactly the loss every classifier already minimises, and mutual information puts a hard ceiling on everything downstream of a sensor. Three results, each unusually sharp.

MathematicsMachine Learning
4 min readProbability Foundations

Which Wrong Distribution Do You Want?

One bimodal target, one Gaussian, and two directions of the same divergence. Minimising KL(P||Q) puts the Gaussian across both modes with almost no mass where the target actually lives; minimising KL(Q||P) puts it on one mode at a value of 0.6931 nats, which is ln 2 to four decimals and not a coincidence. Each fit is judged catastrophic by the other objective, 2.0976 against 15.2799.

Machine LearningMathematics
4 min readStatistical Learning Theory

The Theorem That Says Nothing About Your Problem

Averaged over all 256 functions from three bits to one, a nearest-neighbour learner and a learner built to be wrong on purpose both score exactly 0.500000 off the training set. That is the no free lunch theorem, it is exactly true, and the moment the average is restricted to the six functions that depend on a single bit the two separate to 0.333333 and 0.666667.

Machine LearningMathematics
← Back to all articles