A linear layer can contain millions of weights, yet its backward pass fits on a few lines. Matrix calculus lets us derive those lines with their shapes visible, then compare them with numerical derivatives.
Spotted in the wild
- “d L”First-order loss change.
- “gradient with respect to W”A matrix with the same shape as W.
- “trace of A”Sum of diagonal entries.
- “outer product of g and x”Each entry is .
- “Kronecker delta”One when the indices agree; otherwise zero.
- “diagonal matrix of p”Matrix with p on its diagonal.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “d L” | First-order loss change. | ||
| “gradient with respect to W” | A matrix with the same shape as W. | ||
| “trace of A” | Sum of diagonal entries. | ||
| “outer product of g and x” | Each entry is . | ||
| “Kronecker delta” | One when the indices agree; otherwise zero. | ||
| “diagonal matrix of p” | Matrix with p on its diagonal. |
Fix a convention first
For scalar and column , we write . The gradient has the same shape as . The Jacobian has output coordinates as rows; some texts call this numerator layout. A row derivative of a scalar is the transpose of our column gradient. Read the convention before copying a formula.
For a matrix , the gradient has the shape of and
This trace is just the sum of the entrywise products of the gradient and the perturbation.
If has shape , what is the shape of ?
Quadratic forms and least squares
For , differentiating each occurrence of gives
Only when is symmetric may we simplify this to . The antisymmetric part contributes zero to the scalar quadratic form.
Let and . Then , so . Averaging over examples adds a factor . Omitting the one-half adds a factor 2. A gradient check often catches a lost reduction factor.
Backward through an affine layer
For and upstream column gradient ,
If is , then is and is . Their outer product has exactly the weight shape. A batch sums or averages these outer products according to the definition of its loss.
If and , what is entry (1,2) of using one-based indices?
Softmax and cross-entropy
For logits , set . The quotient rule gives
Adding a constant to every logit leaves softmax unchanged. Correspondingly, each row of sums to zero. In code, subtract the largest logit before exponentiating to avoid overflow.
For targets with ,
Indeed, the th component is . This holds for both one-hot and soft target distributions. If targets are not normalised, the general expression is .
For and , what is the second logit gradient?
Verify shapes and values
Perturb one parameter at a time and compare its central difference with the corresponding gradient entry. Keep the data and any random masks fixed. Avoid activation kinks and try several step sizes. For a large model, a directional check compares with , requiring just two evaluations per direction.
Read beyond
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 5.2–5.4: vector and matrix derivatives
Distinguish a scalar gradient from a vector-valued Jacobian.
Book · free online · ~20 min
Convex OptimizationBoyd & Vandenberghe · Appendix A: mathematical background
Review gradients and Hessians of quadratic functions.
Book · free online · ~20 min
The Matrix Calculus You Need For Deep LearningParr & Howard · The gradient of a neural network loss
Rebuild the example from individual partial derivatives.
Read the equation in context
The Matrix Calculus You Need For Deep LearningTerence Parr & Jeremy Howard · 2018Parr and Howard build neural-network derivatives from scalar rules and Jacobians. The affine map shown here is a compact application of those rules. Expand one output in components, then recover the matrix expression and check its dimensions.
Decode the paper · Vector sum reduction and vector chain rules, written for an affine layer
The Matrix Calculus You Need For Deep LearningTerence Parr & Jeremy Howard · 2018
Parr and Howard build neural-network derivatives from scalar rules and Jacobians. The affine map shown here is a compact application of those rules. Expand one output in components, then recover the matrix expression and check its dimensions.
Options
Your turn
Select a candidate gradient, then compare every component with central differences. A small error supports the formula at the tested point; try more than one point before trusting a general rule.
Interactive lab
Gradient detective
{
"a": [
1.5,
-2
],
"x": [
2,
-1,
0.5
],
"shape": [
2,
3
]
}
Parameter (2 × 3, row order): [2, 1.1, 1.6, 0.3, -0.6, -1.7]Match · Expression ↔ Meaning
Affine gradients
Options
Match · Expression ↔ Meaning
Loss reductions
Options
Proof puzzle
Quadratic gradient
Claim
Derive .
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Softmax ignores a common shift
Claim
Prove and explain why .
Your typeset proof appears here.
Coding problems
Problem 16·Warm-up
Least-squares gradient
Let , , , and . Report the sum of the components of .
Problem 17·Standard
Softmax curvature trace
Logits are . Compute the trace of the softmax Jacobian as a reduced fraction.
Problem 18·Challenge
A batch of outer products
For , a layer sees and upstream gradient . The batch loss is a sum. Report the squared Frobenius norm of .
Key takeaways
- Check gradient shapes before checking numerical values.
- A nonsymmetric quadratic form uses A plus its transpose.
- Affine weight gradients are outer products.
- Softmax cross-entropy gives prediction minus target when targets sum to one.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
For an arbitrary square , what is ?
If softmax outputs , what is ?
What does adding the same constant to every logit do?
For , what is ?
What must stay fixed during a numerical gradient check?
Which expression backpropagates through to x?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Gradient detective (+40 XP)
- Bonus: Problem 16: Least-squares gradient (+20 XP)
- Bonus: Problem 17: Softmax curvature trace (+35 XP)
- Bonus: Problem 18: A batch of outer products (+50 XP)
- Bonus: Proof: Quadratic gradient (+25 XP)
- Bonus: Proof: Softmax ignores a common shift (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Affine gradients (+20 XP)
- Bonus: Match: Loss reductions (+20 XP)