Skip to content
AriadneTechnology

The Middle Ring · Chamber 6 of 9

Matrix Calculus: Gradients of Layers and Losses

Differentiate with respect to vectors and matrices: least squares, linear layers and softmax cross-entropy, checked numerically.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Keep shapes straight with numerator and denominator layout
  • Derive the gradients of linear, quadratic and least-squares functions
  • Derive the gradients of a linear layer and of softmax cross-entropy
  • Verify any gradient with a numerical gradient check
DiscoverLearnRead beyondPapers & lecturesYour turn

A linear layer can contain millions of weights, yet its backward pass fits on a few lines. Matrix calculus lets us derive those lines with their shapes visible, then compare them with numerical derivatives.

Spotted in the wild

∂y∂x=W,y=Wx+b\frac{\partial\mathbf y}{\partial\mathbf x}=W,\qquad \mathbf y=W\mathbf x+\mathbf b
The Matrix Calculus You Need For Deep Learning
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • dLdL“d L”
    First-order loss change.
  • ∇WL\nabla_WL“gradient with respect to W”
    A matrix with the same shape as W.
  • tr⁡(A)\operatorname{tr}(A)“trace of A”
    Sum of diagonal entries.
  • gxTgx^T“outer product of g and x”
    Each entry is gixjg_i x_j.
  • δij\delta_{ij}“Kronecker delta”
    One when the indices agree; otherwise zero.
  • diag⁡(p)\operatorname{diag}(p)“diagonal matrix of p”
    Matrix with p on its diagonal.

Fix a convention first

For scalar L(x)L(x) and column xx, we write dL=(∇xL)TdxdL=(\nabla_xL)^Tdx. The gradient has the same shape as xx. The Jacobian ∂y/∂x\partial y/\partial x has output coordinates as rows; some texts call this numerator layout. A row derivative of a scalar is the transpose of our column gradient. Read the convention before copying a formula.

For a matrix WW, the gradient has the shape of WW and

dL=tr⁡((∇WL)TdW).dL=\operatorname{tr}((\nabla_WL)^T dW).

This trace is just the sum of the entrywise products of the gradient and the perturbation.

Quick check +20 XP

If WW has shape 3×53\times5, what is the shape of ∇WL\nabla_WL?

Quadratic forms and least squares

For L(x)=xTAxL(x)=x^TAx, differentiating each occurrence of xx gives

∇xL=(A+AT)x.\nabla_xL=(A+A^T)x.

Only when AA is symmetric may we simplify this to 2Ax2Ax. The antisymmetric part contributes zero to the scalar quadratic form.

Let r=Xw−yr=Xw-y and L=12∥r∥2L=\tfrac12\|r\|^2. Then dL=rTdr=rTX dwdL=r^Tdr=r^TX\,dw, so ∇wL=XT(Xw−y)\nabla_wL=X^T(Xw-y). Averaging over nn examples adds a factor 1/n1/n. Omitting the one-half adds a factor 2. A gradient check often catches a lost reduction factor.

Backward through an affine layer

For z=Wx+bz=Wx+b and upstream column gradient g=∇zLg=\nabla_zL,

∇xL=WTg,∇WL=gxT,∇bL=g.\nabla_xL=W^Tg,\qquad \nabla_WL=gx^T,\qquad \nabla_bL=g.

If WW is m×nm\times n, then gg is m×1m\times1 and xTx^T is 1×n1\times n. Their outer product has exactly the weight shape. A batch sums or averages these outer products according to the definition of its loss.

Quick check +20 XP

If g=(2,−1)g=(2,-1) and x=(3,4)x=(3,4), what is entry (1,2) of gxTgx^T using one-based indices?

Softmax and cross-entropy

For logits zz, set pi=ezi/∑jezjp_i=e^{z_i}/\sum_j e^{z_j}. The quotient rule gives

∂pi∂zj=pi(δij−pj),J=diag⁡(p)−ppT.\frac{\partial p_i}{\partial z_j}=p_i(\delta_{ij}-p_j),\qquad J=\operatorname{diag}(p)-pp^T.

Adding a constant to every logit leaves softmax unchanged. Correspondingly, each row of JJ sums to zero. In code, subtract the largest logit before exponentiating to avoid overflow.

For targets yi≥0y_i\ge0 with ∑iyi=1\sum_i y_i=1,

L=−∑iyilog⁡pi,∇zL=p−y.L=-\sum_i y_i\log p_i,\qquad \nabla_zL=p-y.

Indeed, the jjth component is −∑iyi(δij−pj)=−yj+pj-\sum_i y_i(\delta_{ij}-p_j)=-y_j+p_j. This holds for both one-hot and soft target distributions. If targets are not normalised, the general expression is p∑iyi−yp\sum_i y_i-y.

Quick check +20 XP

For p=(0.2,0.3,0.5)p=(0.2,0.3,0.5) and y=(0,1,0)y=(0,1,0), what is the second logit gradient?

Verify shapes and values

Perturb one parameter at a time and compare its central difference with the corresponding gradient entry. Keep the data and any random masks fixed. Avoid activation kinks and try several step sizes. For a large model, a directional check compares [L(θ+hv)−L(θ−hv)]/(2h)[L(\theta+hv)-L(\theta-hv)]/(2h) with gTvg^Tv, requiring just two evaluations per direction.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 5.2–5.4: vector and matrix derivatives

Distinguish a scalar gradient from a vector-valued Jacobian.

Book · free online · ~20 min

Convex Optimization

Boyd & Vandenberghe · Appendix A: mathematical background

Review gradients and Hessians of quadratic functions.

Book · free online · ~20 min

The Matrix Calculus You Need For Deep Learning

Parr & Howard · The gradient of a neural network loss

Rebuild the example from individual partial derivatives.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

The Matrix Calculus You Need For Deep LearningTerence Parr & Jeremy Howard · 2018

Parr and Howard build neural-network derivatives from scalar rules and Jacobians. The affine map shown here is a compact application of those rules. Expand one output in components, then recover the matrix expression and check its dimensions.

Decode the paper · Vector sum reduction and vector chain rules, written for an affine layer

The Matrix Calculus You Need For Deep Learning

Terence Parr & Jeremy Howard · 2018

+25 XP
∂y∂x=W,y=Wx+b\frac{\partial\mathbf y}{\partial\mathbf x}=W,\qquad \mathbf y=W\mathbf x+\mathbf b

Parr and Howard build neural-network derivatives from scalar rules and Jacobians. The affine map shown here is a compact application of those rules. Expand one output in components, then recover the matrix expression and check its dimensions.

x\mathbf x
WW
b\mathbf b
∂y/∂x\partial\mathbf y/\partial\mathbf x

Options

The SoftMax Derivative, Step-by-StepStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Select a candidate gradient, then compare every component with central differences. A small error supports the formula at the tested point; try more than one point before trusting a general rule.

Interactive lab

Gradient detective

Choose a formula, then run a central-difference check. Pass all four cases. The fixed data and current parameter vector are shown below.
f(W)=a⊤Wxf(W) = \mathbf{a}^\top W \mathbf{x}
{
  "a": [
    1.5,
    -2
  ],
  "x": [
    2,
    -1,
    0.5
  ],
  "shape": [
    2,
    3
  ]
}
Parameter (2 × 3, row order): [2, 1.1, 1.6, 0.3, -0.6, -1.7]

Challenge: Gradient detectivePick the true gradient of four functions and confirm each one with a gradient check.+40 XP

Match · Expression ↔ Meaning

Affine gradients

+20 XP
∇xL\nabla_x L
∇WL\nabla_W L
∇bL\nabla_b L

Options

Match · Expression ↔ Meaning

Loss reductions

+20 XP
12∑ri2\tfrac12\sum r_i^2
∑ri2\sum r_i^2
1n∑ri2\tfrac1n\sum r_i^2

Options

Proof puzzle

Quadratic gradient

+25 XP

Claim

Derive ∇x(xTAx)=(A+AT)x\nabla_x(x^TAx)=(A+A^T)x.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Softmax ignores a common shift

+35 XP

Claim

Prove softmax⁡(z+c1)=softmax⁡(z)\operatorname{softmax}(z+c\mathbf1)=\operatorname{softmax}(z) and explain why J1=0J\mathbf1=0.

Preview

Your typeset proof appears here.

Coding problems

Problem 16·Warm-up

Least-squares gradient

+20 XP

Let X=[[1,2],[3,4]]X=[[1,2],[3,4]], w=(1,−1)Tw=(1,-1)^T, y=(0,1)Ty=(0,1)^T, and L=12∥Xw−y∥2L=\tfrac12\|Xw-y\|^2. Report the sum of the components of ∇wL\nabla_w L.

An exact integer (or a fraction like 7/12)

Problem 17·Standard

Softmax curvature trace

+35 XP

Logits are (log⁡1,log⁡2,log⁡3)(\log 1,\log 2,\log 3). Compute the trace of the softmax Jacobian as a reduced fraction.

An exact integer (or a fraction like 7/12)

Problem 18·Challenge

A batch of outer products

+50 XP

For k=1,…,100k=1,\ldots,100, a layer sees xk=(k,1)Tx_k=(k,1)^T and upstream gradient gk=(1,−k)Tg_k=(1,-k)^T. The batch loss is a sum. Report the squared Frobenius norm of ∑kgkxkT\sum_k g_kx_k^T.

An exact integer (or a fraction like 7/12)

Key takeaways

  • Check gradient shapes before checking numerical values.
  • A nonsymmetric quadratic form uses A plus its transpose.
  • Affine weight gradients are outer products.
  • Softmax cross-entropy gives prediction minus target when targets sum to one.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

For an arbitrary square AA, what is ∇x(xTAx)\nabla_x(x^TAx)?

Question 2 of 6 +20 XP

If softmax outputs (0.25,0.75)(0.25,0.75), what is ∂p1/∂z2\partial p_1/\partial z_2?

Question 3 of 6 +20 XP

What does adding the same constant to every logit do?

Question 4 of 6 +20 XP

For L(w)=12(2w−3)2L(w)=\tfrac12(2w-3)^2, what is L′(1)L'(1)?

Question 5 of 6 +20 XP

What must stay fixed during a numerical gradient check?

Question 6 of 6 +20 XP

Which expression backpropagates through z=Wx+bz=Wx+b to x?

End of the chamber

Clear this chamber

+60 XPDerivative ShapesQuadratic Form GradientSoftmax JacobianCross-Entropy Gradient