Skip to content
AriadneTechnology

The Outer Ring · Chamber 3 of 9

Expectation, Variance and Initialisation

Variance of sums, covariance and correlation, Markov and Chebyshev, and the variance argument behind Xavier and Kaiming initialisation.

40 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Compute the variance of sums and products of random variables
  • Measure dependence with covariance and correlation, and know their limits
  • Prove Markov's and Chebyshev's inequalities
  • Derive Xavier and Kaiming initialisation from a variance argument
DiscoverLearnRead beyondPapers & lecturesYour turn

Before a network learns anything, its initial weights decide whether signals stay measurable through its layers. A variance calculation supplies a principled scale, including the factor of two used for ReLU networks.

Spotted in the wild

12nlVar⁡(wl)=1\frac12 n_l\operatorname{Var}(w_l)=1
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • E[X]\mathbb E[X]“expectation of X”
    Probability-weighted average.
  • E[X2]\mathbb E[X^2]“second moment of X”
    Average square without subtracting the mean.
  • Var⁡(X)\operatorname{Var}(X)“variance of X”
    Expected squared deviation from the mean.
  • Cov⁡(X,Y)\operatorname{Cov}(X,Y)“covariance of X and Y”
    Expected product of centred deviations.
  • ρXY\rho_{XY}“correlation of X and Y”
    Covariance divided by both positive standard deviations.
  • ninn_{in}“fan in”
    Number of inputs contributing to a unit.

Expectation adds, even with dependence

For discrete X, E[f(X)]=∑xf(x)P(X=x)\mathbb E[f(X)]=\sum_x f(x)P(X=x); for a density, replace the sum with an integral. Linearity gives E[aX+bY]=aEX+bEY\mathbb E[aX+bY]=a\mathbb EX+b\mathbb EY whenever these expectations exist. Independence is not required.

Variance centres the variable: Var⁡(X)=E[(X−μ)2]=E[X2]−μ2\operatorname{Var}(X)=\mathbb E[(X-\mu)^2]=\mathbb E[X^2]-\mu^2. It obeys Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX+b)=a^2\operatorname{Var}(X).

Quick check +20 XP

X has mean 3 and second moment 13. What is its variance?

Covariance records cross terms

Define Cov⁡(X,Y)=E[(X−μX)(Y−μY)]\operatorname{Cov}(X,Y)=\mathbb E[(X-\mu_X)(Y-\mu_Y)]. Expanding a square gives

Var⁡(∑iXi)=∑iVar⁡(Xi)+2∑i<jCov⁡(Xi,Xj).\operatorname{Var}\Big(\sum_iX_i\Big)=\sum_i\operatorname{Var}(X_i)+2\sum_{i<j}\operatorname{Cov}(X_i,X_j).

Independent variables with finite variance have zero covariance. The converse is false: for symmetric X, X and X² can be uncorrelated while clearly dependent. Correlation divides covariance by the two standard deviations, when both are positive, and lies between -1 and 1 by Cauchy–Schwarz.

For independent X and Y, E[XY]=EXEY\mathbb E[XY]=\mathbb EX\mathbb EY and

Var⁡(XY)=Var⁡(X)Var⁡(Y)+Var⁡(X)μY2+Var⁡(Y)μX2.\operatorname{Var}(XY)=\operatorname{Var}(X)\operatorname{Var}(Y)+\operatorname{Var}(X)\mu_Y^2+\operatorname{Var}(Y)\mu_X^2.

Independence matters here because we factor expectations of products.

Quick check +20 XP

X and Y each have variance 1 and covariance 0.5. What is Var(X+Y)?

Bounds that need little information

If X≥0, then X≥a1X≥aX\ge a\mathbf1_{X\ge a} for a>0. Taking expectations gives Markov’s inequality, P(X≥a)≤EX/aP(X\ge a)\le\mathbb EX/a.

Apply it to (X−μ)2(X-\mu)^2 to obtain Chebyshev’s inequality:

P(∣X−μ∣≥t)≤Var⁡(X)t2.P(|X-\mu|\ge t)\le\frac{\operatorname{Var}(X)}{t^2}.

These bounds can be loose, but they do not require Gaussian data. A bound above one conveys no additional information; probabilities already cannot exceed one.

Derive an initialisation scale

Let z=∑i=1nwiaiz=\sum_{i=1}^n w_i a_i, with independent zero-mean weights, independent of the inputs. Cross terms vanish under these assumptions, giving Var⁡(z)=nVar⁡(w)E[a2]\operatorname{Var}(z)=n\operatorname{Var}(w)\mathbb E[a^2] when the input second moments agree.

For small tanh inputs, tanh⁡z≈z\tanh z\approx z, so a fan-in scale Var⁡(w)=1/n\operatorname{Var}(w)=1/n approximately preserves forward variance. Xavier initialisation balances forward and backward requirements with 2/(nin+nout)2/(n_{in}+n_{out}); those requirements coincide when widths agree.

For symmetric zero-mean z, ReLU has E[ReLU⁡(z)2]=12E[z2]\mathbb E[\operatorname{ReLU}(z)^2]=\tfrac12\mathbb E[z^2]. It has a positive mean, so its variance is smaller than this second moment. Compensating the second-moment loss in the next pre-activation gives Kaiming initialisation: Var⁡(w)=2/n\operatorname{Var}(w)=2/n.

For uniform weights on [-a,a], variance is a2/3a^2/3, giving a=6/na=\sqrt{6/n} for ReLU. These arguments describe an initial approximation, not a promise that every trained layer or correlated dataset preserves variance.

Quick check +20 XP

A ReLU layer has fan-in 200. What Kaiming weight variance does the forward argument suggest?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapters 3–5: expectation, variance and covariance

Expand a squared sum and identify every covariance term.

Book · free online · ~20 min

Introduction to Probability

Blitzstein & Hwang · Expectation, moments and inequalities

Prove an inequality by comparing a variable with an indicator.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 6.4–6.5: moments and Gaussian distributions

Keep second moments separate from centred covariance.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet ClassificationKaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2015

He and colleagues derive a weight scale by tracking forward signal second moments and backward gradients under independence and symmetry assumptions. ReLU halves the second moment of a symmetric pre-activation; it does not generally halve the centred variance, because its output has a positive mean.

Decode the paper · Section 2.2: variance-preserving initialisation for rectifiers

Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification

Kaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2015

+25 XP
12nlVar⁡(wl)=1\frac12 n_l\operatorname{Var}(w_l)=1

He and colleagues derive a weight scale by tracking forward signal second moments and backward gradients under independence and symmetry assumptions. ReLU halves the second moment of a symmetric pre-activation; it does not generally halve the centred variance, because its output has a positive mean.

nln_l
Var⁡(wl)\operatorname{Var}(w_l)
12\frac12

Options

Calculating the Mean, Variance and Standard DeviationStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Adjust the weight scale and watch the signal through thirty layers. Compare a small-signal tanh approximation with the ReLU second-moment calculation, then check both against the stated assumptions.

Interactive lab

Signal propagation

Each layer has 128 inputs. Weight variance is scale/128. Starting pre-activation variance is 0.01. We propagate second moments assuming independent zero-mean weights and Gaussian pre-activations; real trained networks may differ.
Signal second moment through thirty layers, relative to the starting value0-9.057.5-6.5415-4.0222.5-1.51301Layerlog₁₀ relative second moment
● Predicted signal● Half of starting signal● Twice starting signal

Weight variance

0.003906

Final / initial second moment

8.957e-10

ReLU halves the second moment of a symmetric input. Tanh is approximately linear only near zero.

Challenge: Signal keeperChoose weight variances that keep activations steady through 30 layers, for tanh and for ReLU.+40 XP

Match · Expression ↔ Meaning

Moment formulas

+20 XP
E[X2]−μ2\mathbb E[X^2]-\mu^2
E[XY]−μXμY\mathbb E[XY]-\mu_X\mu_Y
Cov⁡(X,Y)/(σXσY)\operatorname{Cov}(X,Y)/(\sigma_X\sigma_Y)

Options

Match · Expression ↔ Meaning

Initialisation scales

+20 XP
1/nin1/n_{in}
2/nin2/n_{in}
2/(nin+nout)2/(n_{in}+n_{out})

Options

Proof puzzle

Chebyshev from Markov

+25 XP

Claim

Prove P(∣X−μ∣≥t)≤σ2/t2P(|X-\mu|\ge t)\le\sigma^2/t^2 for finite variance and t>0.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

ReLU and second moments

+35 XP

Claim

For a symmetric variable Z with finite second moment, prove E[max⁡(0,Z)2]=12E[Z2]\mathbb E[\max(0,Z)^2]=\tfrac12\mathbb E[Z^2].

Preview

Your typeset proof appears here.

Coding problems

Problem 7·Warm-up

Variance of correlated eyes

+20 XP

Ten estimates each have variance 4. Every distinct pair has covariance 1. What is the variance of their average? Give one decimal place.

A number, rounded to 1 decimal place

Problem 8·Standard

Uniform initialisation

+35 XP

A ReLU layer has 600 inputs. Choose independent uniform weights on [-a,a] with variance 2/600. Give a to 4 decimal places.

A number, rounded to 4 decimal places

Problem 9·Challenge

How much shared noise remains?

+50 XP

One hundred unbiased estimates each have variance 1 and every pair has correlation 0.1. How many independent estimates with variance 1 would be needed to have variance at most that of their average? Return the smallest integer.

An exact integer (or a fraction like 7/12)

Key takeaways

  • Linearity of expectation survives dependence.
  • Variance of a sum includes covariance terms.
  • Markov and Chebyshev give general bounds with limited assumptions.
  • Initialisation scales follow from second moments and explicit independence assumptions.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

Does zero covariance imply independence?

Question 2 of 6 +20 XP

If Var(X)=4, what is Var(3X+7)?

Question 3 of 6 +20 XP

X is nonnegative with mean 2. What Markov upper bound applies to P(X≥10)?

Question 4 of 6 +20 XP

What Chebyshev bound applies to deviation of at least three standard deviations?

Question 5 of 6 +20 XP

What does ReLU halve under symmetric zero-mean pre-activations?

Question 6 of 6 +20 XP

Which identity requires no independence?

End of the chamber

Clear this chamber

+60 XPVariance of SumsCorrelationChebyshev’s InequalityKaiming Initialisation