Diffusion models gradually add Gaussian noise to data. A useful fact makes the whole forward chain easy to sample: independent Gaussian contributions remain Gaussian, with covariances that add after scaling.
Spotted in the wild
- “mean vector”Coordinatewise expected value.
- “covariance matrix”Expected outer product of centred deviations.
- “Cholesky factorisation”A lower-triangular factor for positive definite covariance.
- “squared Mahalanobis distance”Distance adjusted for covariance geometry.
- “multivariate normal”A jointly Gaussian vector with the specified moments.
- “alpha bar t”Product of signal-retention factors through time t.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “mean vector” | Coordinatewise expected value. | ||
| “covariance matrix” | Expected outer product of centred deviations. | ||
| “Cholesky factorisation” | A lower-triangular factor for positive definite covariance. | ||
| “squared Mahalanobis distance” | Distance adjusted for covariance geometry. | ||
| “multivariate normal” | A jointly Gaussian vector with the specified moments. | ||
| “alpha bar t” | Product of signal-retention factors through time t. |
A covariance matrix is a geometry
For a vector X with mean μ, define . Its diagonal entries are variances; its off-diagonal entries are covariances. It is symmetric and positive semidefinite because .
A nonsingular d-dimensional normal has density
Positive definiteness is needed for this ordinary full-dimensional density. A singular covariance can still define a Gaussian supported on a lower-dimensional plane, but the inverse-and-determinant formula above does not apply.
Why is a covariance matrix positive semidefinite?
Distance measured in standard deviations
The squared Mahalanobis distance is . When , it becomes squared Euclidean distance. For diagonal covariance, it divides each squared coordinate deviation by its variance.
Covariance eigenvectors give ellipse directions, and square roots of eigenvalues give their relative axis lengths. The determinant controls the volume scale. Correlation rotates and stretches the contours; it cannot be chosen independently of a valid positive semidefinite covariance.
For a bivariate Gaussian, zero covariance does imply independence. This is a special property of joint Gaussian distributions, not a general consequence of uncorrelatedness.
Affine transformations and sampling
If X is Gaussian, then is Gaussian with mean and covariance . For sampling, find a Cholesky factor , draw independent standard normal coordinates ε, and set .
For , a factor is . It produces and ; the shared first noise term creates covariance 2.
For , what is entry (1,2) of ?
For a bivariate Gaussian with correlation ρ and positive marginal standard deviations, conditioning on gives
The mean shifts according to the observed coordinate and the uncertainty contracts. This formula assumes a nonsingular joint Gaussian; degenerate cases need separate treatment.
Collapse a forward diffusion chain
Write with independent standard Gaussian noise. Substituting earlier steps repeatedly gives a clean-input coefficient , where .
Conditional noise covariance follows with . Induction gives . Thus a single independent Gaussian draw samples any forward time directly. This algebra is why the opening density holds.
If , what is the standard deviation of forward noise conditional on x0?
Read beyond
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Section 6.5: the multivariate Gaussian
Follow the covariance through an affine transformation.
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapter 6: random vectors
Distinguish marginal and conditional distributions.
Book · free online · ~20 min
Introduction to ProbabilityBlitzstein & Hwang · Multivariate distributions and conditioning
Work through a joint normal example and inspect how correlation changes conditional uncertainty.
Read the equation in context
Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain & Pieter Abbeel · 2020DDPM gives a closed-form distribution for a noisy state conditional on a clean input. The input need not itself be Gaussian. Conditioning on x0 makes the mean fixed; if x0 is drawn from a complex data distribution, the marginal distribution of xt can remain a mixture.
Decode the paper · Equation (4): forward marginal
Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain & Pieter Abbeel · 2020
DDPM gives a closed-form distribution for a noisy state conditional on a clean input. The input need not itself be Gaussian. Conditioning on x0 makes the mean fixed; if x0 is drawn from a complex data distribution, the marginal distribution of xt can remain a mixture.
Options
Your turn
Fit the mean, spreads and correlation of three point clouds. Then inspect conditional slices and watch the forward diffusion process reduce the signal while adding noise.
Interactive lab
Gaussian sculptor
You start from the standard normal N(0, I): the unit circle. Drag the gold centre onto the cloud, then shape the ellipse with the sliders.
Fit to this cloud
Match below 0.05 (the tick). Lower is closer.
Your Gaussian
, major axis at 0°,
Match · Expression ↔ Meaning
Gaussian geometry
Options
Match · Expression ↔ Meaning
Transforms
Options
Proof puzzle
Affine covariance
Claim
If Y=AX+b, prove Cov(Y)=A Cov(X) A transpose.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Diffusion variance by induction
Claim
If v0=0 and , show .
Your typeset proof appears here.
Coding problems
Problem 10·Warm-up
Shape an ellipse
Let Σ=[[4,2],[2,3]]. Compute its determinant.
Problem 11·Standard
Two diffusion steps
Let x0=2, alpha1=0.8 and alpha2=0.45. What is the conditional mean of x2? Give one decimal place.
Problem 12·Challenge
A scheduled noise budget
For t=1,…,100, set alpha_t=t/(t+1). Starting with x0 fixed, report the conditional noise variance after step 100 as a reduced fraction.
Key takeaways
- Covariance determines Gaussian contour shape and orientation.
- Affine maps transform covariance on both sides.
- Cholesky sampling turns independent noise into correlated coordinates.
- Diffusion’s closed form is a conditional Gaussian statement.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
For μ=0, Σ=diag(4,1), what is the squared Mahalanobis distance of x=(2,1)?
What happens if covariance is singular?
A bivariate normal has sigma2=2 and correlation 0.5. What is Var(X2 | X1)?
When does zero covariance guarantee independence?
Two independent zero-mean scalar Gaussians have variances 2 and 3. What is the variance of their sum?
Is the unconditional distribution of noisy data necessarily Gaussian at finite diffusion time?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Gaussian sculptor (+40 XP)
- Bonus: Problem 10: Shape an ellipse (+20 XP)
- Bonus: Problem 11: Two diffusion steps (+35 XP)
- Bonus: Problem 12: A scheduled noise budget (+50 XP)
- Bonus: Proof: Affine covariance (+25 XP)
- Bonus: Proof: Diffusion variance by induction (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Gaussian geometry (+20 XP)
- Bonus: Match: Transforms (+20 XP)