A minibatch gradient is an average of noisy per-example gradients. Larger batches reduce sampling noise, but with diminishing returns. Mean and covariance calculations explain that tradeoff before we appeal to a bell curve.
Spotted in the wild
- “sample mean”Arithmetic average of n observations.
- “standard error of the mean”Sampling spread of an iid average.
- “converges in probability”Fixed-error probabilities tend to zero.
- “converges in distribution”Distribution functions converge at continuity points of the limit.
- “standardised average”Centred mean divided by its standard error.
- “gradient covariance divided by batch size”Covariance of an independent batch average.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “sample mean” | Arithmetic average of n observations. | ||
| “standard error of the mean” | Sampling spread of an iid average. | ||
| “converges in probability” | Fixed-error probabilities tend to zero. | ||
| “converges in distribution” | Distribution functions converge at continuity points of the limit. | ||
| “standardised average” | Centred mean divided by its standard error. | ||
| “gradient covariance divided by batch size” | Covariance of an independent batch average. |
Start with the exact mean and variance
For independent identically distributed X_i with mean μ and finite variance σ², let . Linearity and vanishing cross-covariances give
The standard error is . It describes fluctuations of the estimator across repeated samples. The population standard deviation σ describes the spread of individual observations. Four times as many observations halve the standard error.
Individual observations have standard deviation 6. What is the standard error of the mean of 100 iid observations?
The weak law of large numbers
Apply Chebyshev to the average:
For every fixed positive error tolerance, the probability of exceeding it tends to zero. That is convergence in probability. It does not promise monotonic improvement after every new sample, nor exact equality at a finite n. More general laws relax these assumptions; this proof establishes the finite-variance iid version.
What does the weak law guarantee?
The central limit theorem
Under the same iid assumptions with ,
The unscaled average concentrates at μ. The centred, rescaled error approaches a nondegenerate normal distribution. These are compatible statements about different quantities.
The theorem supplies an asymptotic shape, not an exact finite-sample guarantee. Strong skewness or rare events can require much larger samples than a familiar rule such as “thirty is enough”. A Cauchy population has neither a finite mean nor variance; its sample average remains Cauchy, so this theorem does not apply. Dependence also changes the variance and may require a different limit theorem.
Minibatches and Monte Carlo
For independently sampled per-example gradients g_i with mean G and covariance Σ, a batch average has covariance Σ/B. This is an exact covariance identity; normality is unnecessary. Sampling B examples without replacement from a finite dataset of N reduces covariance further by a finite-population factor. Define ; then the batch covariance is for N>1.
Correlated observations do not give the usual benefit. If scalar errors have common variance σ² and pairwise correlation ρ, the average variance is , when such a joint covariance is valid. The shared component remains as n grows.
Monte Carlo estimates obey the same averaging rules. Report an uncertainty estimate alongside the numerical answer, and preserve independence at the correct unit: repeated frames from the same subject are not equivalent to new subjects.
Multiply an independent minibatch size by 4. By what factor does the gradient standard deviation change?
Read beyond
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapter 7: limit theorems
Separate convergence of averages from their limiting fluctuation shape.
Book · free online · ~20 min
Introduction to ProbabilityBlitzstein & Hwang · Laws of large numbers and central limit theorem
Track which quantity is centred and rescaled.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 6.4 and 8.2: moments and empirical risk
Interpret a sampled objective and its gradient as averages.
Read the equation in context
An Empirical Model of Large-Batch TrainingSam McCandlish, Jared Kaplan, Dario Amodei & OpenAI Dota Team · 2018The paper studies batch sizes through gradient noise and training efficiency. This ratio compares total gradient variance with squared mean-gradient size. It is undefined when the mean gradient is zero and does not by itself prescribe a universally optimal batch size.
Decode the paper · Simple gradient noise scale, with G the population gradient
An Empirical Model of Large-Batch TrainingSam McCandlish, Jared Kaplan, Dario Amodei & OpenAI Dota Team · 2018
The paper studies batch sizes through gradient noise and training efficiency. This ratio compares total gradient variance with squared mean-gradient size. It is undefined when the mean gradient is zero and does not by itself prescribe a universally optimal batch size.
Options
Your turn
Compare the distribution of individual draws with the distribution of standardised averages. Increase the batch size for three skewed populations and inspect both the histogram and the remaining skewness.
Interactive lab
Bell maker
Theoretical skewness of the mean
2.0000
Standard error before standardisation
1.0000
Match · Expression ↔ Meaning
Three different spreads
Options
Match · Expression ↔ Meaning
Assumptions and conclusions
Options
Proof puzzle
Weak law from Chebyshev
Claim
Prove convergence in probability of an iid finite-variance sample mean.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Variance of an average
Claim
For independent X_i with common variance σ², prove Var(mean)=σ²/n.
Your typeset proof appears here.
Coding problems
Problem 13·Warm-up
Halve the error bar
A mean from 64 independent observations has standard error 0.25. How many observations are required for standard error at most 0.1, assuming unchanged population variance?
Problem 14·Standard
A distribution-free guarantee
For iid data with variance 4, Chebyshev should bound P(|mean−μ|≥0.1) by at most 0.05. Find the smallest integer n that satisfies the bound.
Problem 15·Challenge
Exact finite-population variance
A dataset contains values 1,…,100. Draw two distinct values uniformly without replacement and average them. Compute the variance of that average to 4 decimal places.
Key takeaways
- The standard error describes an estimator, not individual observations.
- The weak law controls fixed-error probabilities.
- The CLT describes centred, rescaled fluctuations and needs assumptions.
- Dependence and finite-population sampling change the usual batch-size calculation.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
What becomes approximately standard normal in the iid CLT?
Var(X)=9 for iid observations. What is Var(mean of 36 observations)?
Why is the ordinary iid finite-variance CLT inapplicable to Cauchy observations?
With variance 1 and n=100, what Chebyshev upper bound applies to mean error at least 0.2?
Does the exact covariance-of-an-average formula require Gaussian data?
A large dataset consists of repeated measurements on a few subjects. What needs attention?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Bell maker (+40 XP)
- Bonus: Problem 13: Halve the error bar (+20 XP)
- Bonus: Problem 14: A distribution-free guarantee (+35 XP)
- Bonus: Problem 15: Exact finite-population variance (+50 XP)
- Bonus: Proof: Weak law from Chebyshev (+25 XP)
- Bonus: Proof: Variance of an average (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Three different spreads (+20 XP)
- Bonus: Match: Assumptions and conclusions (+20 XP)