Information Theory and Cross Entropy
Understand entropy, cross entropy, KL divergence and perplexity from first principles
A taste of a lesson
If the true label is class 2 and my model gives it probability 0.25, what is the cross entropy loss?
With a one hot label, cross entropy reduces to minus the log of the probability assigned to the correct class. So the loss is -ln(0.25), about 1.386 nats, or exactly 2 bits in base 2. Intuitively, the model was as uncertain as choosing among four equally likely options. If it had given 0.9, the loss would be about 0.105 nats; at 0.01, about 4.6 nats, which shows how harshly confident mistakes are punished. Exercise: what probability on the correct class gives a loss of exactly 1 bit?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Compute surprise and entropy in bits and nats
- Define cross entropy and explain it as a coding cost
- Compute KL divergence and explain its asymmetry
- Show why minimising cross entropy equals maximum likelihood
- Interpret perplexity and mutual information correctly
Lesson plan
- 1 Surprise and information Quantify surprise as minus log probability. Start
- 2 Entropy Compute entropy as average surprise. Start
- 3 Cross entropy Define cross entropy and interpret it as a mismatch cost. Start
- 4 KL divergence Measure the extra cost of using q instead of p. Start
- 5 Training with cross entropy Connect cross entropy to maximum likelihood and perplexity. Start
- 6 Mutual information and smoothing Use mutual information and label smoothing appropriately. Start
Try asking
About this tutor
An advanced tutor for learners who use cross entropy and want to know what it really measures. You will build information theory from the idea of surprise, compute entropy for coins and dice, then define cross entropy and KL divergence and see why minimising cross entropy trains classifiers and language models. Lessons cover the asymmetry of KL divergence, bits versus nats, perplexity as an evaluation measure, mutual information and practical topics such as label smoothing. Every quantity is computed on small distributions so you can check it by hand.
Reviews
Students can review a tutor after a paid lesson. Nobody has yet.
About the teacher
Probability for machine learning, plus forecasting and anomaly detection
9 tutors 361 lessons taught Sample
I teach probability the way machine learning uses it: random variables, likelihood, entropy and simulation. I also teach two applied areas where probability matters every day: time series forecasting and anomaly detection. My work background is in monitoring and forecasting for operational systems, where wrong alarms and missed incidents both have a cost. I teach through small simulations, coin and...
See Kenta's profile and tutorsMore like this
Other tutors on the same or nearby topics.