Before a network learns anything, its initial weights decide whether signals stay measurable through its layers. A variance calculation supplies a principled scale, including the factor of two used for ReLU networks.
Spotted in the wild
- “expectation of X”Probability-weighted average.
- “second moment of X”Average square without subtracting the mean.
- “variance of X”Expected squared deviation from the mean.
- “covariance of X and Y”Expected product of centred deviations.
- “correlation of X and Y”Covariance divided by both positive standard deviations.
- “fan in”Number of inputs contributing to a unit.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “expectation of X” | Probability-weighted average. | ||
| “second moment of X” | Average square without subtracting the mean. | ||
| “variance of X” | Expected squared deviation from the mean. | ||
| “covariance of X and Y” | Expected product of centred deviations. | ||
| “correlation of X and Y” | Covariance divided by both positive standard deviations. | ||
| “fan in” | Number of inputs contributing to a unit. |
Expectation adds, even with dependence
For discrete X, ; for a density, replace the sum with an integral. Linearity gives whenever these expectations exist. Independence is not required.
Variance centres the variable: . It obeys .
X has mean 3 and second moment 13. What is its variance?
Covariance records cross terms
Define . Expanding a square gives
Independent variables with finite variance have zero covariance. The converse is false: for symmetric X, X and X² can be uncorrelated while clearly dependent. Correlation divides covariance by the two standard deviations, when both are positive, and lies between -1 and 1 by Cauchy–Schwarz.
For independent X and Y, and
Independence matters here because we factor expectations of products.
X and Y each have variance 1 and covariance 0.5. What is Var(X+Y)?
Bounds that need little information
If X≥0, then for a>0. Taking expectations gives Markov’s inequality, .
Apply it to to obtain Chebyshev’s inequality:
These bounds can be loose, but they do not require Gaussian data. A bound above one conveys no additional information; probabilities already cannot exceed one.
Derive an initialisation scale
Let , with independent zero-mean weights, independent of the inputs. Cross terms vanish under these assumptions, giving when the input second moments agree.
For small tanh inputs, , so a fan-in scale approximately preserves forward variance. Xavier initialisation balances forward and backward requirements with ; those requirements coincide when widths agree.
For symmetric zero-mean z, ReLU has . It has a positive mean, so its variance is smaller than this second moment. Compensating the second-moment loss in the next pre-activation gives Kaiming initialisation: .
For uniform weights on [-a,a], variance is , giving for ReLU. These arguments describe an initial approximation, not a promise that every trained layer or correlated dataset preserves variance.
A ReLU layer has fan-in 200. What Kaiming weight variance does the forward argument suggest?
Read beyond
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapters 3–5: expectation, variance and covariance
Expand a squared sum and identify every covariance term.
Book · free online · ~20 min
Introduction to ProbabilityBlitzstein & Hwang · Expectation, moments and inequalities
Prove an inequality by comparing a variable with an indicator.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 6.4–6.5: moments and Gaussian distributions
Keep second moments separate from centred covariance.
Read the equation in context
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet ClassificationKaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2015He and colleagues derive a weight scale by tracking forward signal second moments and backward gradients under independence and symmetry assumptions. ReLU halves the second moment of a symmetric pre-activation; it does not generally halve the centred variance, because its output has a positive mean.
Decode the paper · Section 2.2: variance-preserving initialisation for rectifiers
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet ClassificationKaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2015
He and colleagues derive a weight scale by tracking forward signal second moments and backward gradients under independence and symmetry assumptions. ReLU halves the second moment of a symmetric pre-activation; it does not generally halve the centred variance, because its output has a positive mean.
Options
Your turn
Adjust the weight scale and watch the signal through thirty layers. Compare a small-signal tanh approximation with the ReLU second-moment calculation, then check both against the stated assumptions.
Interactive lab
Signal propagation
Weight variance
0.003906
Final / initial second moment
8.957e-10
ReLU halves the second moment of a symmetric input. Tanh is approximately linear only near zero.
Match · Expression ↔ Meaning
Moment formulas
Options
Match · Expression ↔ Meaning
Initialisation scales
Options
Proof puzzle
Chebyshev from Markov
Claim
Prove for finite variance and t>0.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
ReLU and second moments
Claim
For a symmetric variable Z with finite second moment, prove .
Your typeset proof appears here.
Coding problems
Problem 7·Warm-up
Variance of correlated eyes
Ten estimates each have variance 4. Every distinct pair has covariance 1. What is the variance of their average? Give one decimal place.
Problem 8·Standard
Uniform initialisation
A ReLU layer has 600 inputs. Choose independent uniform weights on [-a,a] with variance 2/600. Give a to 4 decimal places.
Problem 9·Challenge
How much shared noise remains?
One hundred unbiased estimates each have variance 1 and every pair has correlation 0.1. How many independent estimates with variance 1 would be needed to have variance at most that of their average? Return the smallest integer.
Key takeaways
- Linearity of expectation survives dependence.
- Variance of a sum includes covariance terms.
- Markov and Chebyshev give general bounds with limited assumptions.
- Initialisation scales follow from second moments and explicit independence assumptions.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Does zero covariance imply independence?
If Var(X)=4, what is Var(3X+7)?
X is nonnegative with mean 2. What Markov upper bound applies to P(X≥10)?
What Chebyshev bound applies to deviation of at least three standard deviations?
What does ReLU halve under symmetric zero-mean pre-activations?
Which identity requires no independence?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Signal keeper (+40 XP)
- Bonus: Problem 7: Variance of correlated eyes (+20 XP)
- Bonus: Problem 8: Uniform initialisation (+35 XP)
- Bonus: Problem 9: How much shared noise remains? (+50 XP)
- Bonus: Proof: Chebyshev from Markov (+25 XP)
- Bonus: Proof: ReLU and second moments (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Moment formulas (+20 XP)
- Bonus: Match: Initialisation scales (+20 XP)