Skip to content
AriadneTechnology

The Inner Ring · Chamber 7 of 9

Entropy, Cross-Entropy and KL Divergence

Measure surprise in bits: entropy, cross-entropy, KL divergence, mutual information and the evidence lower bound.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Compute entropy, cross-entropy and KL divergence, in bits and in nats
  • Prove that the KL divergence is never negative
  • Interpret the cross-entropy loss and perplexity of a language model
  • Derive the evidence lower bound with Jensen's inequality
DiscoverLearnRead beyondPapers & lecturesYour turn

Cross-entropy measures the surprise a model assigns to observed outcomes. Variational inference goes further: it bounds a difficult log-evidence calculation using an expectation and a KL divergence. Both begin with the same logarithm of probability.

Spotted in the wild

L(q)=Eq[log⁡pθ(x∣z)]−DKL(q(z)∥p(z))\mathcal L(q)=\mathbb E_q[\log p_\theta(x\mid z)]-D_{KL}(q(z)\|p(z))
Auto-Encoding Variational Bayes
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • −log⁡p(x)-\log p(x)“surprise of x”
    Information associated with an outcome probability.
  • H(P)H(P)“entropy of P”
    Expected surprise under the same distribution.
  • H(P,Q)H(P,Q)“cross-entropy”
    Expected model surprise when observations follow P.
  • DKL(P∥Q)D_{KL}(P\|Q)“KL divergence from P to Q”
    Expected log probability ratio under P.
  • I(X;Y)I(X;Y)“mutual information”
    Dependence measured against the product of marginals.
  • L(q)\mathcal L(q)“evidence lower bound”
    A tractable lower bound on log model evidence.

Surprise and average surprise

An event of probability p carries surprise −log⁡p-\log p. Independent event probabilities multiply, so their surprises add. Base-two logs measure bits; natural logs measure nats.

For a discrete distribution P,

H(P)=−∑xp(x)log⁡p(x).H(P)=-\sum_xp(x)\log p(x).

Use the limit convention 0log⁡0=00\log0=0. A certain outcome has zero entropy. A uniform distribution on K outcomes has entropy log K, the maximum on that support. A fair coin has one bit; four equally likely outcomes have two bits.

Quick check +20 XP

What is the entropy in bits of four equally likely outcomes?

Cross-entropy and KL

If data follows P but a model predicts Q, its expected surprise is H(P,Q)=−∑xp(x)log⁡q(x)H(P,Q)=-\sum_xp(x)\log q(x). Subtracting entropy gives

DKL(P∥Q)=∑xp(x)log⁡p(x)q(x)=H(P,Q)−H(P).D_{KL}(P\|Q)=\sum_xp(x)\log\frac{p(x)}{q(x)}=H(P,Q)-H(P).

When p(x)>0 but q(x)=0, both KL and cross-entropy are infinite. Zero p terms contribute zero. KL is generally asymmetric and is not a metric: it does not obey all distance axioms.

To prove nonnegativity where the ratios are defined, use log⁡u≤u−1\log u\le u-1 with u=q/pu=q/p. Then −∑plog⁡(q/p)≥∑(p−q)-\sum p\log(q/p)\ge\sum(p-q) over the support of P, which is nonnegative. Equality requires the distributions to agree.

Minimising cross-entropy over Q therefore minimises KL from the fixed data distribution. Entropy of P is constant during that optimisation.

Quick check +20 XP

If P gives an outcome positive mass and Q gives it zero mass, what is KL(P || Q)?

Perplexity and shared information

Perplexity is the exponential of mean negative log probability: exp⁡(H)\exp(H) for nats or 2H2^H for bits. A uniform model over K symbols has perplexity K. Compare language-model perplexities only when tokenisation, dataset and averaging conventions match.

Mutual information is I(X;Y)=DKL(PXY∥PXPY)I(X;Y)=D_{KL}(P_{XY}\|P_XP_Y): the divergence between the joint distribution and the distribution that would hold under independence. It is nonnegative and vanishes exactly at independence. It can detect dependence even when covariance is zero.

Differential entropy for continuous variables uses a density inside an integral. Unlike discrete entropy, it can be negative and depends on coordinate scale. Do not transfer every discrete intuition to continuous entropy.

Derive the evidence lower bound

Introduce a latent variable z and an approximate posterior q with appropriate support. Using Jensen’s inequality for concave log,

log⁡pθ(x)=log⁡Eq[pθ(x,z)q(z)]≥Eq[log⁡pθ(x,z)−log⁡q(z)].\log p_\theta(x)=\log\mathbb E_q\left[\frac{p_\theta(x,z)}{q(z)}\right]\ge\mathbb E_q[\log p_\theta(x,z)-\log q(z)].

Split the joint into likelihood times prior to recover the opening ELBO. Equivalently,

log⁡pθ(x)=L(q)+DKL(q(z)∥pθ(z∣x)).\log p_\theta(x)=\mathcal L(q)+D_{KL}(q(z)\|p_\theta(z\mid x)).

The rightmost divergence is the gap. Optimising the bound changes both the generative model and its approximate inference distribution; it is not generally the same as evaluating an exact log-likelihood.

Quick check +20 XP

When is the ELBO equal to log evidence?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Deep Learning

Goodfellow, Bengio & Courville · Section 3.13: information theory

Work out the entropy and KL divergence of a tiny categorical distribution.

Book · free online · ~20 min

Dive into Deep Learning

Zhang, Lipton, Li & Smola · Section 22.11: information theory

Track whether logarithms use base two or base e.

Book · free online · ~20 min

Convex Optimization

Boyd & Vandenberghe · Section 3.1: convex functions and Jensen’s inequality

Apply the concavity of log to an expectation.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Auto-Encoding Variational BayesDiederik P. Kingma & Max Welling · 2014

The paper optimises a lower bound on log evidence with an approximate posterior. The first term rewards explaining observations, and the second measures divergence from the prior. The bound equals the log evidence only when the approximate posterior matches the model’s exact posterior.

Decode the paper · Equation (3): evidence lower bound, writing q for the approximate posterior

Auto-Encoding Variational Bayes

Diederik P. Kingma & Max Welling · 2014

+25 XP
L(q)=Eq[log⁡pθ(x∣z)]−DKL(q(z)∥p(z))\mathcal L(q)=\mathbb E_q[\log p_\theta(x\mid z)]-D_{KL}(q(z)\|p(z))

The paper optimises a lower bound on log evidence with an approximate posterior. The first term rewards explaining observations, and the second measures divergence from the prior. The bound equals the log evidence only when the approximate posterior matches the model’s exact posterior.

q(z)q(z)
pθ(x∣z)p_\theta(x\mid z)
p(z)p(z)
DKLD_{KL}

Options

Entropy (for data science), Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Tune a binary distribution to three entropy targets. Then compare KL in both directions for two distinct distributions.

Interactive lab

Entropy sculptor

Shape P=(p,1−p) to match three entropy targets, then compare it with Q=(q,1−q). Entropy is in bits; the two KL divergences use nats.
Binary entropy as a function of success probability000.250.2750.50.550.750.82511.1pEntropy in bits
● Binary entropy● Target● Current entropy

H(P), bits

0.88129

KL(P ∥ Q), nats

0.33892

KL(Q ∥ P), nats

0.33892

Challenge: Entropy sculptorShape distributions to hit three entropy targets, then make two KL divergences disagree.+40 XP

Match · Expression ↔ Meaning

Information measures

+20 XP
−log⁡p(x)-\log p(x)
EP[−log⁡q(X)]\mathbb E_P[-\log q(X)]
DKL(PXY∥PXPY)D_{KL}(P_{XY}\|P_XP_Y)

Options

Match · Expression ↔ Meaning

Units and bounds

+20 XP
log⁡2\log_2
log⁡e\log_e
log⁡p(x)−L\log p(x)-\mathcal L

Options

Proof puzzle

Cross-entropy decomposition

+25 XP

Claim

Show H(P,Q)=H(P)+DKL(P∥Q)H(P,Q)=H(P)+D_{KL}(P\|Q) for finite values.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Uniform has maximum entropy

+35 XP

Claim

Prove H(P)≤log K for a distribution on K outcomes.

Preview

Your typeset proof appears here.

Coding problems

Problem 19·Warm-up

Entropy of a small alphabet

+20 XP

For P=(1/2,1/4,1/4), compute entropy in bits to 4 decimal places.

A number, rounded to 4 decimal places

Problem 20·Standard

Asymmetric divergence

+35 XP

P=(1/2,1/2), Q=(1/4,3/4). Compute KL(P || Q) in nats to 6 decimal places.

A number, rounded to 6 decimal places

Problem 21·Challenge

A latent-variable bound

+50 XP

A model has joint probabilities p(x,z=0)=0.2 and p(x,z=1)=0.3 for the observed x. Choose q(z)=(0.5,0.5). Compute log p(x) minus the ELBO in nats to 6 decimal places.

A number, rounded to 6 decimal places

Key takeaways

  • State the log base and the distribution defining the expectation.
  • Cross-entropy equals data entropy plus KL divergence.
  • Perplexity comparisons require matching evaluation conventions.
  • The ELBO is a bound with a specific posterior-approximation gap.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

A model has mean negative log probability of 3 bits per token. What is its perplexity?

Question 2 of 6 +20 XP

Cross-entropy is 2.5 nats and data entropy is 1.7 nats. What is KL?

Question 3 of 6 +20 XP

Is KL generally symmetric?

Question 4 of 6 +20 XP

Which quantity is constant when fitting Q to a fixed P?

Question 5 of 6 +20 XP

A certain discrete outcome has what entropy?

Question 6 of 6 +20 XP

What does zero mutual information imply?

End of the chamber

Clear this chamber

+60 XPEntropyPerplexityKL DivergenceEvidence Lower Bound