Cross-entropy measures the surprise a model assigns to observed outcomes. Variational inference goes further: it bounds a difficult log-evidence calculation using an expectation and a KL divergence. Both begin with the same logarithm of probability.
Spotted in the wild
- “surprise of x”Information associated with an outcome probability.
- “entropy of P”Expected surprise under the same distribution.
- “cross-entropy”Expected model surprise when observations follow P.
- “KL divergence from P to Q”Expected log probability ratio under P.
- “mutual information”Dependence measured against the product of marginals.
- “evidence lower bound”A tractable lower bound on log model evidence.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “surprise of x” | Information associated with an outcome probability. | ||
| “entropy of P” | Expected surprise under the same distribution. | ||
| “cross-entropy” | Expected model surprise when observations follow P. | ||
| “KL divergence from P to Q” | Expected log probability ratio under P. | ||
| “mutual information” | Dependence measured against the product of marginals. | ||
| “evidence lower bound” | A tractable lower bound on log model evidence. |
Surprise and average surprise
An event of probability p carries surprise . Independent event probabilities multiply, so their surprises add. Base-two logs measure bits; natural logs measure nats.
For a discrete distribution P,
Use the limit convention . A certain outcome has zero entropy. A uniform distribution on K outcomes has entropy log K, the maximum on that support. A fair coin has one bit; four equally likely outcomes have two bits.
What is the entropy in bits of four equally likely outcomes?
Cross-entropy and KL
If data follows P but a model predicts Q, its expected surprise is . Subtracting entropy gives
When p(x)>0 but q(x)=0, both KL and cross-entropy are infinite. Zero p terms contribute zero. KL is generally asymmetric and is not a metric: it does not obey all distance axioms.
To prove nonnegativity where the ratios are defined, use with . Then over the support of P, which is nonnegative. Equality requires the distributions to agree.
Minimising cross-entropy over Q therefore minimises KL from the fixed data distribution. Entropy of P is constant during that optimisation.
If P gives an outcome positive mass and Q gives it zero mass, what is KL(P || Q)?
Perplexity and shared information
Perplexity is the exponential of mean negative log probability: for nats or for bits. A uniform model over K symbols has perplexity K. Compare language-model perplexities only when tokenisation, dataset and averaging conventions match.
Mutual information is : the divergence between the joint distribution and the distribution that would hold under independence. It is nonnegative and vanishes exactly at independence. It can detect dependence even when covariance is zero.
Differential entropy for continuous variables uses a density inside an integral. Unlike discrete entropy, it can be negative and depends on coordinate scale. Do not transfer every discrete intuition to continuous entropy.
Derive the evidence lower bound
Introduce a latent variable z and an approximate posterior q with appropriate support. Using Jensen’s inequality for concave log,
Split the joint into likelihood times prior to recover the opening ELBO. Equivalently,
The rightmost divergence is the gap. Optimising the bound changes both the generative model and its approximate inference distribution; it is not generally the same as evaluating an exact log-likelihood.
When is the ELBO equal to log evidence?
Read beyond
Book · free online · ~20 min
Deep LearningGoodfellow, Bengio & Courville · Section 3.13: information theory
Work out the entropy and KL divergence of a tiny categorical distribution.
Book · free online · ~20 min
Dive into Deep LearningZhang, Lipton, Li & Smola · Section 22.11: information theory
Track whether logarithms use base two or base e.
Book · free online · ~20 min
Convex OptimizationBoyd & Vandenberghe · Section 3.1: convex functions and Jensen’s inequality
Apply the concavity of log to an expectation.
Read the equation in context
Auto-Encoding Variational BayesDiederik P. Kingma & Max Welling · 2014The paper optimises a lower bound on log evidence with an approximate posterior. The first term rewards explaining observations, and the second measures divergence from the prior. The bound equals the log evidence only when the approximate posterior matches the model’s exact posterior.
Decode the paper · Equation (3): evidence lower bound, writing q for the approximate posterior
Auto-Encoding Variational BayesDiederik P. Kingma & Max Welling · 2014
The paper optimises a lower bound on log evidence with an approximate posterior. The first term rewards explaining observations, and the second measures divergence from the prior. The bound equals the log evidence only when the approximate posterior matches the model’s exact posterior.
Options
Your turn
Tune a binary distribution to three entropy targets. Then compare KL in both directions for two distinct distributions.
Interactive lab
Entropy sculptor
H(P), bits
0.88129
KL(P ∥ Q), nats
0.33892
KL(Q ∥ P), nats
0.33892
Match · Expression ↔ Meaning
Information measures
Options
Match · Expression ↔ Meaning
Units and bounds
Options
Proof puzzle
Cross-entropy decomposition
Claim
Show for finite values.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Uniform has maximum entropy
Claim
Prove H(P)≤log K for a distribution on K outcomes.
Your typeset proof appears here.
Coding problems
Problem 19·Warm-up
Entropy of a small alphabet
For P=(1/2,1/4,1/4), compute entropy in bits to 4 decimal places.
Problem 20·Standard
Asymmetric divergence
P=(1/2,1/2), Q=(1/4,3/4). Compute KL(P || Q) in nats to 6 decimal places.
Problem 21·Challenge
A latent-variable bound
A model has joint probabilities p(x,z=0)=0.2 and p(x,z=1)=0.3 for the observed x. Choose q(z)=(0.5,0.5). Compute log p(x) minus the ELBO in nats to 6 decimal places.
Key takeaways
- State the log base and the distribution defining the expectation.
- Cross-entropy equals data entropy plus KL divergence.
- Perplexity comparisons require matching evaluation conventions.
- The ELBO is a bound with a specific posterior-approximation gap.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
A model has mean negative log probability of 3 bits per token. What is its perplexity?
Cross-entropy is 2.5 nats and data entropy is 1.7 nats. What is KL?
Is KL generally symmetric?
Which quantity is constant when fitting Q to a fixed P?
A certain discrete outcome has what entropy?
What does zero mutual information imply?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Entropy sculptor (+40 XP)
- Bonus: Problem 19: Entropy of a small alphabet (+20 XP)
- Bonus: Problem 20: Asymmetric divergence (+35 XP)
- Bonus: Problem 21: A latent-variable bound (+50 XP)
- Bonus: Proof: Cross-entropy decomposition (+25 XP)
- Bonus: Proof: Uniform has maximum entropy (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Information measures (+20 XP)
- Bonus: Match: Units and bounds (+20 XP)