A language model assigns conditional probabilities to the next word. Training can choose parameters that make the observed sequence more probable. The same principle yields least squares for Gaussian observations and cross-entropy for classification.
Spotted in the wild
- “likelihood of theta given fixed data”The observation-model expression viewed as a parameter function.
- “log-likelihood”Logarithm of the likelihood.
- “theta hat”An estimator computed from data.
- “parameter maximising”Returns a parameter rather than the objective value.
- “bias of an estimator”Expected estimate minus the true parameter.
- “negative log-likelihood”A loss formed by negating log probability or density.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “likelihood of theta given fixed data” | The observation-model expression viewed as a parameter function. | ||
| “log-likelihood” | Logarithm of the likelihood. | ||
| “theta hat” | An estimator computed from data. | ||
| “parameter maximising” | Returns a parameter rather than the objective value. | ||
| “bias of an estimator” | Expected estimate minus the true parameter. | ||
| “negative log-likelihood” | A loss formed by negating log probability or density. |
Reverse what is fixed
A probability model varies x while θ is fixed. A likelihood uses the same expression with the observed x fixed and θ varying. It need not integrate to one as a function of θ. A continuous-data likelihood is a density value, not the probability of one exact observation.
For iid observations, . Maximising it is equivalent to maximising because log is strictly increasing. Working with logs also avoids multiplying many tiny values.
In a likelihood calculation, what is held fixed?
Derive three estimators
For s successes in n Bernoulli trials,
Differentiate to get . If s=0 or s=n, the maximum is at the boundary and must be checked separately; an interior derivative equation is insufficient.
For exponential waiting times with rate λ, , giving when the total waiting time is positive.
For normal data with unknown mean and variance, differentiating gives
A zero empirical variance creates a degenerate boundary case if the model requires strictly positive variance.
Six successes occur in eight Bernoulli trials. What is the MLE of p?
Likelihood and bias are different questions
An estimator is unbiased if its expectation over repeated datasets equals the true parameter. An MLE need not have this property.
The identity
shows why the normal MLE variance is biased. For iid finite-variance data, taking expectations gives . Thus division by n has expectation , and division by n−1 gives an unbiased variance estimator when n>1. Unbiasedness alone does not guarantee lower mean squared estimation error.
Familiar losses from observation models
Suppose with independent Gaussian errors of fixed variance σ². The negative log-likelihood is a constant plus . With σ fixed, maximising likelihood is equivalent to minimising squared error.
For a categorical label y, the negative log-likelihood is , equal to one-hot cross-entropy. For soft target distributions, cross-entropy is an expected categorical negative log-likelihood.
These equivalences depend on the noise model. Heavy-tailed errors may motivate a different likelihood. A likelihood can be maximised accurately even when the model is misspecified; the fit does not certify that its assumptions match the data.
Squared error is proportional to a Gaussian negative log-likelihood under which condition?
Read beyond
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Chapter 8: when models meet data
Derive a loss by writing down an observation model.
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapter 8: statistical inference
Compare estimator definitions with properties such as bias.
Book · free online · ~20 min
Deep LearningGoodfellow, Bengio & Courville · Section 5.5: maximum likelihood estimation
Keep the observed data fixed while varying the parameter.
Read the equation in context
A Neural Probabilistic Language ModelYoshua Bengio, Réjean Ducharme, Pascal Vincent & Christian Jauvin · 2003The paper learns distributed word representations through a probabilistic language model. The displayed objective keeps its log-likelihood term and explicitly omits regularisation. Sequence factorisation uses conditional probabilities; it does not require adjacent words to be independent.
Decode the paper · Log-likelihood part of the training criterion, omitting regularisation and a constant averaging factor
A Neural Probabilistic Language ModelYoshua Bengio, Réjean Ducharme, Pascal Vincent & Christian Jauvin · 2003
The paper learns distributed word representations through a probabilistic language model. The displayed objective keeps its log-likelihood term and explicitly omits regularisation. Sequence factorisation uses conditional probabilities; it does not require adjacent words to be independent.
Options
Your turn
Move the model parameters while the data stays fixed. Compare each observation’s log-likelihood contribution, then locate the maximum for three datasets.
Interactive lab
Likelihood climber
Observed data: 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1, 0, 1, 1, 0, 1, 0, 1, 1, 0
Sum of log-likelihood contributions
-13.8629
Each observation’s contribution
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
- -0.6931
Match · Expression ↔ Meaning
Observation models
Options
Match · Expression ↔ Meaning
Estimator properties
Options
Proof puzzle
Bernoulli MLE
Claim
For 0<s<n successes, derive p hat = s/n.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Why one degree of freedom is lost
Claim
For iid data with variance σ², show .
Your typeset proof appears here.
Coding problems
Problem 16·Warm-up
Fit a waiting-time model
Observations are for k=1,…,100. Compute the exponential MLE rate to 8 decimal places.
Problem 17·Standard
MLE or unbiased?
For data 1,…,100, compute the difference between the unbiased sample variance and the normal MLE variance. Give 4 decimal places.
Problem 18·Challenge
Compare two probability models
Ten independent Bernoulli observations contain seven successes. Compute the likelihood ratio L(p=0.7)/L(p=0.5) to 6 decimal places.
Key takeaways
- Likelihood compares parameter choices on fixed data.
- Log-likelihood turns independent products into sums.
- Maximum likelihood and unbiasedness are different properties.
- Squared error and cross-entropy follow from explicit observation models.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Exponential waiting times are 1, 2 and 3. What is the MLE rate?
Normal observations are 1 and 3. What is the MLE variance?
For observations 1 and 3, what is the unbiased sample variance?
Must a likelihood integrate to one over parameters?
Is every maximum-likelihood estimator unbiased?
All Bernoulli observations are zero. Where is the MLE of p?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Likelihood climber (+30 XP)
- Bonus: Problem 16: Fit a waiting-time model (+20 XP)
- Bonus: Problem 17: MLE or unbiased? (+35 XP)
- Bonus: Problem 18: Compare two probability models (+50 XP)
- Bonus: Proof: Bernoulli MLE (+25 XP)
- Bonus: Proof: Why one degree of freedom is lost (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Observation models (+20 XP)
- Bonus: Match: Estimator properties (+20 XP)