Information Theory
The one place in this subject where a bound is met exactly: entropy is the shortest any code can be, the best code reaches it, and the surcharge for using the wrong distribution is the loss function you already train with.
Sign in to take quizzes, earn XP, and unlock stages as you reach 90% mastery.
Entropy and the Shortest Code
25 min · 100 XPWhy entropy is a limit rather than a summary, a code that meets it to the last decimal, the source where whole bits are too lumpy to reach it, and the trick that closes the gap.
Open lesson →Sign in to take the 3-question quiz.
The Cost of the Wrong Distribution
30 min · 120 XPCross-entropy as the bill for coding one distribution with another’s code, the surcharge that is exactly the KL divergence, why that surcharge is the loss you already minimise, and what a confident mistake costs.
Open lesson →Sign in to take the 3-question quiz.
Channels, Dependence and What Cannot Be Recovered
30 min · 120 XPWhat a noisy channel can carry per use, a dependence that correlation reports as zero and mutual information reports as a whole bit, and the theorem that says no amount of processing recovers what was destroyed.
Open lesson →Sign in to take the 3-question quiz.