Skip to content
Kudos AI

Entropy and information calculator

Edit two distributions side by side and read off entropy, cross entropy, and KL divergence in bits, with the identity H(p, q) = H(p) + KL(p ‖ q) held on screen.

FreeInformation TheoryProbabilityMathematics

Entropy and information

Everything in bits (base 2).

H(p) entropy
1.0000 bits
H(p, q) cross entropy
1.7370 bits
KL(p ‖ q) divergence
0.7370 bits

H(p, q) = H(p) + KL(p ‖ q) → 1.7370 = 1.0000 + 0.7370

Maximum for 2 outcomes: 1.0000 bits

p - the true distribution

A50.0%
B50.0%

q - the believed distribution

A90.0%
B10.0%

Presets

Entropy is highest when every outcome is equally likely and falls to zero once one outcome is certain. Cross entropy is what a model is trained to minimise; it splits exactly into the entropy of the data - which no model can remove - plus the KL divergence, the part caused by believing the wrong distribution.

Runs entirely in your browser. Nothing you enter is uploaded or stored.

The ideas behind it

10 min readProbability Foundations

Entropy and Information

Measuring uncertainty in bits: Shannon entropy and why the logarithm is base 2, information gain worked on a split, and how cross-entropy and KL divergence relate to entropy and to the loss functions used to train classifiers.

Information TheoryProbabilityMathematics
7 min readSupervised Learning

Decision Trees and Ensembles

How recursive binary splitting builds a tree, why the Gini index beats accuracy as a splitting criterion, and how bagging and random forests turn a high-variance learner into a strong one, with the split arithmetic worked out.

Machine LearningStatistics