Skip to content
AriadneTechnology

The Middle Ring · Chamber 4 of 9

The Multivariate Gaussian

Covariance matrices, ellipses of probability, sampling with Cholesky, and the fact about adding Gaussians that powers diffusion models.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Read and evaluate the multivariate normal density
  • Relate a covariance matrix to the shape and orientation of its ellipses
  • Push Gaussians through affine maps, and sample them with a Cholesky factor
  • Derive the closed-form noising step of a diffusion model
DiscoverLearnRead beyondPapers & lecturesYour turn

Diffusion models gradually add Gaussian noise to data. A useful fact makes the whole forward chain easy to sample: independent Gaussian contributions remain Gaussian, with covariances that add after scaling.

Spotted in the wild

q(xt∣x0)=N ⁣(xt;αˉtx0,(1−αˉt)I)q(x_t\mid x_0)=\mathcal N\!\left(x_t;\sqrt{\bar\alpha_t}x_0,(1-\bar\alpha_t)I\right)
Denoising Diffusion Probabilistic Models
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • μ\mu“mean vector”
    Coordinatewise expected value.
  • Σ\Sigma“covariance matrix”
    Expected outer product of centred deviations.
  • Σ=LLT\Sigma=LL^T“Cholesky factorisation”
    A lower-triangular factor for positive definite covariance.
  • (x−μ)TΣ−1(x−μ)(x-\mu)^T\Sigma^{-1}(x-\mu)“squared Mahalanobis distance”
    Distance adjusted for covariance geometry.
  • N(μ,Σ)\mathcal N(\mu,\Sigma)“multivariate normal”
    A jointly Gaussian vector with the specified moments.
  • αˉt\bar\alpha_t“alpha bar t”
    Product of signal-retention factors through time t.

A covariance matrix is a geometry

For a vector X with mean μ, define Σ=E[(X−μ)(X−μ)T]\Sigma=\mathbb E[(X-\mu)(X-\mu)^T]. Its diagonal entries are variances; its off-diagonal entries are covariances. It is symmetric and positive semidefinite because vTΣv=Var⁡(vTX)≥0v^T\Sigma v=\operatorname{Var}(v^TX)\ge0.

A nonsingular d-dimensional normal has density

p(x)=exp⁡[−12(x−μ)TΣ−1(x−μ)](2π)d/2det⁡Σ.p(x)=\frac{\exp[-\tfrac12(x-\mu)^T\Sigma^{-1}(x-\mu)]}{(2\pi)^{d/2}\sqrt{\det\Sigma}}.

Positive definiteness is needed for this ordinary full-dimensional density. A singular covariance can still define a Gaussian supported on a lower-dimensional plane, but the inverse-and-determinant formula above does not apply.

Quick check +20 XP

Why is a covariance matrix positive semidefinite?

Distance measured in standard deviations

The squared Mahalanobis distance is (x−μ)TΣ−1(x−μ)(x-\mu)^T\Sigma^{-1}(x-\mu). When Σ=I\Sigma=I, it becomes squared Euclidean distance. For diagonal covariance, it divides each squared coordinate deviation by its variance.

Covariance eigenvectors give ellipse directions, and square roots of eigenvalues give their relative axis lengths. The determinant controls the volume scale. Correlation rotates and stretches the contours; it cannot be chosen independently of a valid positive semidefinite covariance.

For a bivariate Gaussian, zero covariance does imply independence. This is a special property of joint Gaussian distributions, not a general consequence of uncorrelatedness.

Affine transformations and sampling

If X is Gaussian, then Y=AX+bY=AX+b is Gaussian with mean Aμ+bA\mu+b and covariance AΣATA\Sigma A^T. For sampling, find a Cholesky factor Σ=LLT\Sigma=LL^T, draw independent standard normal coordinates ε, and set X=μ+LϵX=\mu+L\epsilon.

For Σ=[[4,2],[2,2]]\Sigma=[[4,2],[2,2]], a factor is L=[[2,0],[1,1]]L=[[2,0],[1,1]]. It produces X1=2ϵ1X_1=2\epsilon_1 and X2=ϵ1+ϵ2X_2=\epsilon_1+\epsilon_2; the shared first noise term creates covariance 2.

Quick check +20 XP

For L=[[2,0],[1,1]]L=[[2,0],[1,1]], what is entry (1,2) of LLTLL^T?

For a bivariate Gaussian with correlation ρ and positive marginal standard deviations, conditioning on X1=aX_1=a gives

E[X2∣X1=a]=μ2+ρσ2σ1(a−μ1),Var⁡(X2∣X1=a)=σ22(1−ρ2).\mathbb E[X_2\mid X_1=a]=\mu_2+\rho\frac{\sigma_2}{\sigma_1}(a-\mu_1),\qquad \operatorname{Var}(X_2\mid X_1=a)=\sigma_2^2(1-\rho^2).

The mean shifts according to the observed coordinate and the uncertainty contracts. This formula assumes a nonsingular joint Gaussian; degenerate cases need separate treatment.

Collapse a forward diffusion chain

Write xt=αtxt−1+1−αtϵtx_t=\sqrt{\alpha_t}x_{t-1}+\sqrt{1-\alpha_t}\epsilon_t with independent standard Gaussian noise. Substituting earlier steps repeatedly gives a clean-input coefficient αˉt\sqrt{\bar\alpha_t}, where αˉt=∏s=1tαs\bar\alpha_t=\prod_{s=1}^t\alpha_s.

Conditional noise covariance follows vt=αtvt−1+1−αtv_t=\alpha_t v_{t-1}+1-\alpha_t with v0=0v_0=0. Induction gives vt=1−αˉtv_t=1-\bar\alpha_t. Thus a single independent Gaussian draw samples any forward time directly. This algebra is why the opening density holds.

Quick check +20 XP

If αˉt=0.36\bar\alpha_t=0.36, what is the standard deviation of forward noise conditional on x0?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Section 6.5: the multivariate Gaussian

Follow the covariance through an affine transformation.

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapter 6: random vectors

Distinguish marginal and conditional distributions.

Book · free online · ~20 min

Introduction to Probability

Blitzstein & Hwang · Multivariate distributions and conditioning

Work through a joint normal example and inspect how correlation changes conditional uncertainty.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain & Pieter Abbeel · 2020

DDPM gives a closed-form distribution for a noisy state conditional on a clean input. The input need not itself be Gaussian. Conditioning on x0 makes the mean fixed; if x0 is drawn from a complex data distribution, the marginal distribution of xt can remain a mixture.

Decode the paper · Equation (4): forward marginal

Denoising Diffusion Probabilistic Models

Jonathan Ho, Ajay Jain & Pieter Abbeel · 2020

+25 XP
q(xt∣x0)=N ⁣(xt;αˉtx0,(1−αˉt)I)q(x_t\mid x_0)=\mathcal N\!\left(x_t;\sqrt{\bar\alpha_t}x_0,(1-\bar\alpha_t)I\right)

DDPM gives a closed-form distribution for a noisy state conditional on a clean input. The input need not itself be Gaussian. Conditioning on x0 makes the mean fixed; if x0 is drawn from a complex data distribution, the marginal distribution of xt can remain a mixture.

x0x_0
xtx_t
αˉt\bar\alpha_t
II

Options

Covariance, Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Fit the mean, spreads and correlation of three point clouds. Then inspect conditional slices and watch the forward diffusion process reduce the signal while adding noise.

Interactive lab

Gaussian sculptor

The teal dots are 300 draws from a hidden Gaussian. The gold ellipses are the 1σ and 2σ contours of yours, N(μ,Σ)\mathcal{N}(\boldsymbol{\mu}, \Sigma), with its axes along the eigenvectors of Σ\Sigma. Set a mean, two spreads and a correlation until your Gaussian matches each cloud: KL(cloud ∥ yours)<0.05\mathrm{KL}(\text{cloud} \,\Vert\, \text{yours}) < 0.05.
−4−4−2−2002244x₁x₂
0/3 matched

You start from the standard normal N(0, I): the unit circle. Drag the gold centre onto the cloud, then shape the ellipse with the sliders.

Fit to this cloud

KL\mathrm{KL}2.41 nats

Match below 0.05 (the tick). Lower is closer.

Your Gaussian

Σ=(1.000.000.001.00)\Sigma = \begin{pmatrix} 1.00 & 0.00 \\ 0.00 & 1.00 \end{pmatrix}
L=(1.0000.001.00)L = \begin{pmatrix} 1.00 & 0 \\ 0.00 & 1.00 \end{pmatrix}

λ1=1.00, λ2=1.00\lambda_1 = 1.00,\ \lambda_2 = 1.00, major axis at 0°, det⁡Σ=1.00\det \Sigma = 1.00

Challenge: Gaussian sculptorMatch three clouds of points by setting a mean, two spreads and a correlation.+40 XP

Match · Expression ↔ Meaning

Gaussian geometry

+20 XP
μ\mu
λi(Σ)\lambda_i(\Sigma)
det⁡Σ\sqrt{\det\Sigma}

Options

Match · Expression ↔ Meaning

Transforms

+20 XP
μ+Lϵ\mu+L\epsilon
Aμ+bA\mu+b
AΣATA\Sigma A^T

Options

Proof puzzle

Affine covariance

+25 XP

Claim

If Y=AX+b, prove Cov(Y)=A Cov(X) A transpose.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Diffusion variance by induction

+35 XP

Claim

If v0=0 and vt=αtvt−1+1−αtv_t=\alpha_t v_{t-1}+1-\alpha_t, show vt=1−∏s=1tαsv_t=1-\prod_{s=1}^t\alpha_s.

Preview

Your typeset proof appears here.

Coding problems

Problem 10·Warm-up

Shape an ellipse

+20 XP

Let Σ=[[4,2],[2,3]]. Compute its determinant.

An exact integer (or a fraction like 7/12)

Problem 11·Standard

Two diffusion steps

+35 XP

Let x0=2, alpha1=0.8 and alpha2=0.45. What is the conditional mean of x2? Give one decimal place.

A number, rounded to 1 decimal place

Problem 12·Challenge

A scheduled noise budget

+50 XP

For t=1,…,100, set alpha_t=t/(t+1). Starting with x0 fixed, report the conditional noise variance after step 100 as a reduced fraction.

An exact integer (or a fraction like 7/12)

Key takeaways

  • Covariance determines Gaussian contour shape and orientation.
  • Affine maps transform covariance on both sides.
  • Cholesky sampling turns independent noise into correlated coordinates.
  • Diffusion’s closed form is a conditional Gaussian statement.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

For μ=0, Σ=diag(4,1), what is the squared Mahalanobis distance of x=(2,1)?

Question 2 of 6 +20 XP

What happens if covariance is singular?

Question 3 of 6 +20 XP

A bivariate normal has sigma2=2 and correlation 0.5. What is Var(X2 | X1)?

Question 4 of 6 +20 XP

When does zero covariance guarantee independence?

Question 5 of 6 +20 XP

Two independent zero-mean scalar Gaussians have variances 2 and 3. What is the variance of their sum?

Question 6 of 6 +20 XP

Is the unconditional distribution of noisy data necessarily Gaussian at finite diffusion time?

End of the chamber

Clear this chamber

+60 XPMultivariate GaussianCovariance MatrixMahalanobis DistanceForward Diffusion