Skip to content
AriadneTechnology

The Middle Ring · Chamber 5 of 9

Large Numbers and the Central Limit Theorem

Why averages settle, why they settle into a bell curve, and what that says about minibatches, error bars and Monte Carlo.

40 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Derive the mean and variance of a sample average
  • Prove the weak law of large numbers
  • State the central limit theorem and recognise when it fails
  • Explain how the batch size controls the noise in a minibatch gradient
DiscoverLearnRead beyondPapers & lecturesYour turn

A minibatch gradient is an average of noisy per-example gradients. Larger batches reduce sampling noise, but with diminishing returns. Mean and covariance calculations explain that tradeoff before we appeal to a bell curve.

Spotted in the wild

Bnoise=tr⁡(Σ)∥G∥2\mathcal B_{\mathrm{noise}}=\frac{\operatorname{tr}(\Sigma)}{\|G\|^2}
An Empirical Model of Large-Batch Training
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • Xˉn\bar X_n“sample mean”
    Arithmetic average of n observations.
  • σ/n\sigma/\sqrt n“standard error of the mean”
    Sampling spread of an iid average.
  • →P\xrightarrow{P}“converges in probability”
    Fixed-error probabilities tend to zero.
  • →d\xrightarrow{d}“converges in distribution”
    Distribution functions converge at continuity points of the limit.
  • ZnZ_n“standardised average”
    Centred mean divided by its standard error.
  • Σ/B\Sigma/B“gradient covariance divided by batch size”
    Covariance of an independent batch average.

Start with the exact mean and variance

For independent identically distributed X_i with mean μ and finite variance σ², let Xˉn=n−1∑iXi\bar X_n=n^{-1}\sum_iX_i. Linearity and vanishing cross-covariances give

EXˉn=μ,Var⁡(Xˉn)=σ2/n.\mathbb E\bar X_n=\mu,\qquad \operatorname{Var}(\bar X_n)=\sigma^2/n.

The standard error is σ/n\sigma/\sqrt n. It describes fluctuations of the estimator across repeated samples. The population standard deviation σ describes the spread of individual observations. Four times as many observations halve the standard error.

Quick check +20 XP

Individual observations have standard deviation 6. What is the standard error of the mean of 100 iid observations?

The weak law of large numbers

Apply Chebyshev to the average:

P(∣Xˉn−μ∣≥ε)≤σ2nε2⟶0.P(|\bar X_n-\mu|\ge\varepsilon)\le\frac{\sigma^2}{n\varepsilon^2}\longrightarrow0.

For every fixed positive error tolerance, the probability of exceeding it tends to zero. That is convergence in probability. It does not promise monotonic improvement after every new sample, nor exact equality at a finite n. More general laws relax these assumptions; this proof establishes the finite-variance iid version.

Quick check +20 XP

What does the weak law guarantee?

The central limit theorem

Under the same iid assumptions with 0<σ2<∞0<\sigma^2<\infty,

Zn=n(Xˉn−μ)σ→dN(0,1).Z_n=\frac{\sqrt n(\bar X_n-\mu)}{\sigma}\xrightarrow{d}\mathcal N(0,1).

The unscaled average concentrates at μ. The centred, rescaled error approaches a nondegenerate normal distribution. These are compatible statements about different quantities.

The theorem supplies an asymptotic shape, not an exact finite-sample guarantee. Strong skewness or rare events can require much larger samples than a familiar rule such as “thirty is enough”. A Cauchy population has neither a finite mean nor variance; its sample average remains Cauchy, so this theorem does not apply. Dependence also changes the variance and may require a different limit theorem.

Minibatches and Monte Carlo

For independently sampled per-example gradients g_i with mean G and covariance Σ, a batch average has covariance Σ/B. This is an exact covariance identity; normality is unnecessary. Sampling B examples without replacement from a finite dataset of N reduces covariance further by a finite-population factor. Define Σ=N−1∑i(gi−G)(gi−G)T\Sigma=N^{-1}\sum_i(g_i-G)(g_i-G)^T; then the batch covariance is N−BB(N−1)Σ\frac{N-B}{B(N-1)}\Sigma for N>1.

Correlated observations do not give the usual benefit. If scalar errors have common variance σ² and pairwise correlation ρ, the average variance is σ2[ρ+(1−ρ)/n]\sigma^2[\rho+(1-\rho)/n], when such a joint covariance is valid. The shared component remains as n grows.

Monte Carlo estimates obey the same averaging rules. Report an uncertainty estimate alongside the numerical answer, and preserve independence at the correct unit: repeated frames from the same subject are not equivalent to new subjects.

Quick check +20 XP

Multiply an independent minibatch size by 4. By what factor does the gradient standard deviation change?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapter 7: limit theorems

Separate convergence of averages from their limiting fluctuation shape.

Book · free online · ~20 min

Introduction to Probability

Blitzstein & Hwang · Laws of large numbers and central limit theorem

Track which quantity is centred and rescaled.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 6.4 and 8.2: moments and empirical risk

Interpret a sampled objective and its gradient as averages.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

An Empirical Model of Large-Batch TrainingSam McCandlish, Jared Kaplan, Dario Amodei & OpenAI Dota Team · 2018

The paper studies batch sizes through gradient noise and training efficiency. This ratio compares total gradient variance with squared mean-gradient size. It is undefined when the mean gradient is zero and does not by itself prescribe a universally optimal batch size.

Decode the paper · Simple gradient noise scale, with G the population gradient

An Empirical Model of Large-Batch Training

Sam McCandlish, Jared Kaplan, Dario Amodei & OpenAI Dota Team · 2018

+25 XP
Bnoise=tr⁡(Σ)∥G∥2\mathcal B_{\mathrm{noise}}=\frac{\operatorname{tr}(\Sigma)}{\|G\|^2}

The paper studies batch sizes through gradient noise and training efficiency. This ratio compares total gradient variance with squared mean-gradient size. It is undefined when the mean gradient is zero and does not by itself prescribe a universally optimal batch size.

GG
Σ\Sigma
tr⁡(Σ)\operatorname{tr}(\Sigma)
Bnoise\mathcal B_{\mathrm{noise}}

Options

The Central Limit Theorem, Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Compare the distribution of individual draws with the distribution of standardised averages. Increase the batch size for three skewed populations and inspect both the histogram and the remaining skewness.

Interactive lab

Bell maker

Draw 1,200 independent batches and standardise their averages. For each population, find the smallest integer n with theoretical mean skewness below 0.3. This is a lab criterion, not a universal normal-approximation guarantee.
Histogram of standardised sample means compared with the standard normal density-40-1.50.26610.5323.50.79861.06√n (sample mean − population mean) / population SDDensity
● Batch means (displayed range)● Standard normal

Theoretical skewness of the mean

2.0000

Standard error before standardisation

1.0000

Challenge: Bell makerFor three skewed distributions, find the smallest sample size whose standardised mean has theoretical skewness below 0.3.+40 XP

Match · Expression ↔ Meaning

Three different spreads

+20 XP
σ\sigma
σ/n\sigma/\sqrt n
n(Xˉn−μ)/σ\sqrt n(\bar X_n-\mu)/\sigma

Options

Match · Expression ↔ Meaning

Assumptions and conclusions

+20 XP
Cov⁡(gi,gj)=0\operatorname{Cov}(g_i,g_j)=0 for distinct samples
σ2<∞\sigma^2<\infty
ρ>0\rho>0 shared across estimates

Options

Proof puzzle

Weak law from Chebyshev

+25 XP

Claim

Prove convergence in probability of an iid finite-variance sample mean.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Variance of an average

+35 XP

Claim

For independent X_i with common variance σ², prove Var(mean)=σ²/n.

Preview

Your typeset proof appears here.

Coding problems

Problem 13·Warm-up

Halve the error bar

+20 XP

A mean from 64 independent observations has standard error 0.25. How many observations are required for standard error at most 0.1, assuming unchanged population variance?

An exact integer (or a fraction like 7/12)

Problem 14·Standard

A distribution-free guarantee

+35 XP

For iid data with variance 4, Chebyshev should bound P(|mean−μ|≥0.1) by at most 0.05. Find the smallest integer n that satisfies the bound.

An exact integer (or a fraction like 7/12)

Problem 15·Challenge

Exact finite-population variance

+50 XP

A dataset contains values 1,…,100. Draw two distinct values uniformly without replacement and average them. Compute the variance of that average to 4 decimal places.

A number, rounded to 4 decimal places

Key takeaways

  • The standard error describes an estimator, not individual observations.
  • The weak law controls fixed-error probabilities.
  • The CLT describes centred, rescaled fluctuations and needs assumptions.
  • Dependence and finite-population sampling change the usual batch-size calculation.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

What becomes approximately standard normal in the iid CLT?

Question 2 of 6 +20 XP

Var(X)=9 for iid observations. What is Var(mean of 36 observations)?

Question 3 of 6 +20 XP

Why is the ordinary iid finite-variance CLT inapplicable to Cauchy observations?

Question 4 of 6 +20 XP

With variance 1 and n=100, what Chebyshev upper bound applies to mean error at least 0.2?

Question 5 of 6 +20 XP

Does the exact covariance-of-an-average formula require Gaussian data?

Question 6 of 6 +20 XP

A large dataset consists of repeated measurements on a few subjects. What needs attention?

End of the chamber

Clear this chamber

+60 XPLaw of Large NumbersStandard ErrorCentral Limit TheoremGradient Noise