Skip to content
AriadneTechnology

The Middle Ring · Chamber 6 of 9

Maximum Likelihood: Fitting Models to Data

Choose the parameters that make the data most probable, and find squared error and cross-entropy hiding inside.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Write likelihoods and log-likelihoods for independent data
  • Derive maximum-likelihood estimators with calculus
  • Prove that the maximum-likelihood variance is biased, and correct it
  • Show that squared error and cross-entropy are negative log-likelihoods
DiscoverLearnRead beyondPapers & lecturesYour turn

A language model assigns conditional probabilities to the next word. Training can choose parameters that make the observed sequence more probable. The same principle yields least squares for Gaussian observations and cross-entropy for classification.

Spotted in the wild

ℓ(θ)=∑tlog⁡pθ(wt∣wt−1,…,wt−n+1)\ell(\theta)=\sum_t\log p_\theta(w_t\mid w_{t-1},\ldots,w_{t-n+1})
A Neural Probabilistic Language Model
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • L(θ;x)L(\theta;x)“likelihood of theta given fixed data”
    The observation-model expression viewed as a parameter function.
  • ℓ(θ)\ell(\theta)“log-likelihood”
    Logarithm of the likelihood.
  • θ^\hat\theta“theta hat”
    An estimator computed from data.
  • argmax⁡θ\operatorname*{argmax}_\theta“parameter maximising”
    Returns a parameter rather than the objective value.
  • Bias⁡(θ^)\operatorname{Bias}(\hat\theta)“bias of an estimator”
    Expected estimate minus the true parameter.
  • NLL⁡\operatorname{NLL}“negative log-likelihood”
    A loss formed by negating log probability or density.

Reverse what is fixed

A probability model pθ(x)p_\theta(x) varies x while θ is fixed. A likelihood L(θ;x)L(\theta;x) uses the same expression with the observed x fixed and θ varying. It need not integrate to one as a function of θ. A continuous-data likelihood is a density value, not the probability of one exact observation.

For iid observations, L(θ)=∏ipθ(xi)L(\theta)=\prod_i p_\theta(x_i). Maximising it is equivalent to maximising ℓ(θ)=∑ilog⁡pθ(xi)\ell(\theta)=\sum_i\log p_\theta(x_i) because log is strictly increasing. Working with logs also avoids multiplying many tiny values.

Quick check +20 XP

In a likelihood calculation, what is held fixed?

Derive three estimators

For s successes in n Bernoulli trials,

ℓ(p)=slog⁡p+(n−s)log⁡(1−p),p^=s/n.\ell(p)=s\log p+(n-s)\log(1-p),\qquad \hat p=s/n.

Differentiate to get s/p−(n−s)/(1−p)=0s/p-(n-s)/(1-p)=0. If s=0 or s=n, the maximum is at the boundary and must be checked separately; an interior derivative equation is insufficient.

For exponential waiting times with rate λ, ℓ(λ)=nlog⁡λ−λ∑ixi\ell(\lambda)=n\log\lambda-\lambda\sum_i x_i, giving λ^=n/∑ixi\hat\lambda=n/\sum_i x_i when the total waiting time is positive.

For normal data with unknown mean and variance, differentiating gives

μ^=xˉ,σ^MLE2=1n∑i(xi−xˉ)2.\hat\mu=\bar x,\qquad \hat\sigma^2_{MLE}=\frac1n\sum_i(x_i-\bar x)^2.

A zero empirical variance creates a degenerate boundary case if the model requires strictly positive variance.

Quick check +20 XP

Six successes occur in eight Bernoulli trials. What is the MLE of p?

Likelihood and bias are different questions

An estimator is unbiased if its expectation over repeated datasets equals the true parameter. An MLE need not have this property.

The identity

∑i(Xi−Xˉ)2=∑i(Xi−μ)2−n(Xˉ−μ)2\sum_i(X_i-\bar X)^2=\sum_i(X_i-\mu)^2-n(\bar X-\mu)^2

shows why the normal MLE variance is biased. For iid finite-variance data, taking expectations gives nσ2−n(σ2/n)=(n−1)σ2n\sigma^2-n(\sigma^2/n)=(n-1)\sigma^2. Thus division by n has expectation (n−1)σ2/n(n-1)\sigma^2/n, and division by n−1 gives an unbiased variance estimator when n>1. Unbiasedness alone does not guarantee lower mean squared estimation error.

Familiar losses from observation models

Suppose yi=fθ(xi)+ϵiy_i=f_\theta(x_i)+\epsilon_i with independent Gaussian errors of fixed variance σ². The negative log-likelihood is a constant plus ∑i(yi−fθ(xi))2/(2σ2)\sum_i(y_i-f_\theta(x_i))^2/(2\sigma^2). With σ fixed, maximising likelihood is equivalent to minimising squared error.

For a categorical label y, the negative log-likelihood is −log⁡pθ(y∣x)-\log p_\theta(y\mid x), equal to one-hot cross-entropy. For soft target distributions, cross-entropy is an expected categorical negative log-likelihood.

These equivalences depend on the noise model. Heavy-tailed errors may motivate a different likelihood. A likelihood can be maximised accurately even when the model is misspecified; the fit does not certify that its assumptions match the data.

Quick check +20 XP

Squared error is proportional to a Gaussian negative log-likelihood under which condition?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Chapter 8: when models meet data

Derive a loss by writing down an observation model.

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapter 8: statistical inference

Compare estimator definitions with properties such as bias.

Book · free online · ~20 min

Deep Learning

Goodfellow, Bengio & Courville · Section 5.5: maximum likelihood estimation

Keep the observed data fixed while varying the parameter.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

A Neural Probabilistic Language ModelYoshua Bengio, Réjean Ducharme, Pascal Vincent & Christian Jauvin · 2003

The paper learns distributed word representations through a probabilistic language model. The displayed objective keeps its log-likelihood term and explicitly omits regularisation. Sequence factorisation uses conditional probabilities; it does not require adjacent words to be independent.

Decode the paper · Log-likelihood part of the training criterion, omitting regularisation and a constant averaging factor

A Neural Probabilistic Language Model

Yoshua Bengio, Réjean Ducharme, Pascal Vincent & Christian Jauvin · 2003

+25 XP
ℓ(θ)=∑tlog⁡pθ(wt∣wt−1,…,wt−n+1)\ell(\theta)=\sum_t\log p_\theta(w_t\mid w_{t-1},\ldots,w_{t-n+1})

The paper learns distributed word representations through a probabilistic language model. The displayed objective keeps its log-likelihood term and explicitly omits regularisation. Sequence factorisation uses conditional probabilities; it does not require adjacent words to be independent.

θ\theta
wtw_t
wt−1,…,wt−n+1w_{t-1},\ldots,w_{t-n+1}
ℓ(θ)\ell(\theta)

Options

Maximum Likelihood, Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Move the model parameters while the data stays fixed. Compare each observation’s log-likelihood contribution, then locate the maximum for three datasets.

Interactive lab

Likelihood climber

The observations stay fixed while you move the parameters. Fit a Bernoulli probability, a Gaussian mean and standard deviation, and an exponential rate.

Observed data: 1, 1, 0, 1, 1, 0, 1, 1, 1, 0, 1, 0, 1, 1, 0, 1, 0, 1, 1, 0

Log-likelihood while varying the selected parameter0.01-60.90.255-48.70.5-36.40.745-24.20.99-11.9Success probabilityLog-likelihood
● Likelihood curve● Current parameters

Sum of log-likelihood contributions

-13.8629

Each observation’s contribution
  1. -0.6931
  2. -0.6931
  3. -0.6931
  4. -0.6931
  5. -0.6931
  6. -0.6931
  7. -0.6931
  8. -0.6931
  9. -0.6931
  10. -0.6931
  11. -0.6931
  12. -0.6931
  13. -0.6931
  14. -0.6931
  15. -0.6931
  16. -0.6931
  17. -0.6931
  18. -0.6931
  19. -0.6931
  20. -0.6931

Challenge: Likelihood climberFind the maximum-likelihood estimate for three datasets, each to within 1%.+30 XP

Match · Expression ↔ Meaning

Observation models

+20 XP
Gaussian errors with fixed variance
Categorical labels
Exponential waiting times

Options

Match · Expression ↔ Meaning

Estimator properties

+20 XP
Eθ^=θ\mathbb E\hat\theta=\theta
θ^=arg⁡max⁡θL(θ)\hat\theta=\arg\max_\theta L(\theta)
E[(θ^−θ)2]\mathbb E[(\hat\theta-\theta)^2]

Options

Proof puzzle

Bernoulli MLE

+25 XP

Claim

For 0<s<n successes, derive p hat = s/n.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Why one degree of freedom is lost

+35 XP

Claim

For iid data with variance σ², show E∑i(Xi−Xˉ)2=(n−1)σ2\mathbb E\sum_i(X_i-\bar X)^2=(n-1)\sigma^2.

Preview

Your typeset proof appears here.

Coding problems

Problem 16·Warm-up

Fit a waiting-time model

+20 XP

Observations are xk=k/10x_k=k/10 for k=1,…,100. Compute the exponential MLE rate to 8 decimal places.

A number, rounded to 8 decimal places

Problem 17·Standard

MLE or unbiased?

+35 XP

For data 1,…,100, compute the difference between the unbiased sample variance and the normal MLE variance. Give 4 decimal places.

A number, rounded to 4 decimal places

Problem 18·Challenge

Compare two probability models

+50 XP

Ten independent Bernoulli observations contain seven successes. Compute the likelihood ratio L(p=0.7)/L(p=0.5) to 6 decimal places.

A number, rounded to 6 decimal places

Key takeaways

  • Likelihood compares parameter choices on fixed data.
  • Log-likelihood turns independent products into sums.
  • Maximum likelihood and unbiasedness are different properties.
  • Squared error and cross-entropy follow from explicit observation models.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

Exponential waiting times are 1, 2 and 3. What is the MLE rate?

Question 2 of 6 +20 XP

Normal observations are 1 and 3. What is the MLE variance?

Question 3 of 6 +20 XP

For observations 1 and 3, what is the unbiased sample variance?

Question 4 of 6 +20 XP

Must a likelihood integrate to one over parameters?

Question 5 of 6 +20 XP

Is every maximum-likelihood estimator unbiased?

Question 6 of 6 +20 XP

All Bernoulli observations are zero. Where is the MLE of p?

End of the chamber

Clear this chamber

+60 XPLikelihoodMaximum LikelihoodUnbiased EstimatorLoss as Likelihood