Skip to content
AriadneTechnology

The Middle Ring · Chamber 5 of 9

Jacobians and Automatic Differentiation

The multivariable chain rule is matrix multiplication. Forward mode, reverse mode, and why one backward pass trains a network.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Build the Jacobian of a vector-valued function and read it as a local linear map
  • Compose Jacobians with the multivariable chain rule
  • Compute derivatives with dual numbers (forward mode) and adjoints (reverse mode)
  • Explain why reverse mode is the right choice for training
DiscoverLearnRead beyondPapers & lecturesYour turn

A network maps a vector of parameters to a scalar loss. Writing out every derivative would be wasteful. Automatic differentiation applies local derivative rules to the executed computation and propagates only the products needed for that loss.

Spotted in the wild

vˉi=∑j: vi→vjvˉj∂vj∂vi\bar v_i=\sum_{j:\,v_i\to v_j}\bar v_j\frac{\partial v_j}{\partial v_i}
Automatic Differentiation in Machine Learning: a Survey
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • JFJ_F“Jacobian of F”
    Output-by-input matrix of partial derivatives.
  • JvJv“J times v”
    Forward directional derivative of a vector output.
  • JTvJ^Tv“J transpose times v”
    Backward propagation of an output sensitivity.
  • vˉ\bar v“v bar”
    Adjoint: derivative of the scalar loss with respect to v.
  • ϵ2=0\epsilon^2=0“epsilon squared equals zero”
    The defining dual-number rule; epsilon is formal, not a tiny float.
  • det⁡J\det J“determinant of J”
    Signed local volume scaling for a square Jacobian.

One derivative for every input-output pair

For F:Rn→RmF:\mathbb R^n\to\mathbb R^m, the Jacobian has shape m×nm\times n:

JF(x)ij=∂Fi∂xj,F(x+h)≈F(x)+JF(x)h.J_F(x)_{ij}=\frac{\partial F_i}{\partial x_j},\qquad F(x+h)\approx F(x)+J_F(x)h.

Rows correspond to outputs and columns to inputs. If F(x,y)=(x2+y,xy)F(x,y)=(x^2+y,xy), then

JF(x,y)=[2x1yx].J_F(x,y)=\begin{bmatrix}2x&1\\y&x\end{bmatrix}.

At (1,2)(1,2), a perturbation (0.01,0)(0.01,0) predicts output changes (0.02,0.02)(0.02,0.02). The first output's exact change has an additional 0.00010.0001 from the square.

Quick check +20 XP

For F:R3→R2F:\mathbb R^3\to\mathbb R^2, how many entries are in JFJ_F?

Composition becomes multiplication

For F:Rn→RmF:\mathbb R^n\to\mathbb R^m and G:Rm→RpG:\mathbb R^m\to\mathbb R^p,

JG∘F(x)=JG(F(x))JF(x).J_{G\circ F}(x)=J_G(F(x))J_F(x).

The shapes are (p×m)(m×n)=p×n(p\times m)(m\times n)=p\times n. The order follows the forward composition. For a scalar loss LL after FF, column gradients travel backward as ∇xL=JFT∇FL\nabla_x L=J_F^T\nabla_{F}L.

A square Jacobian's determinant measures signed local area or volume scaling. A zero determinant means the linear approximation loses a dimension. It does not by itself prove global non-invertibility: x↦x3x\mapsto x^3 is one-to-one even though its derivative vanishes at zero.

Quick check +20 XP

If G∘FG\circ F is a composition, which is its Jacobian?

Forward mode: carry a tangent

A dual number a+bϵa+b\epsilon has ϵ2=0\epsilon^2=0. Multiplication gives (a+bϵ)(c+dϵ)=ac+(ad+bc)ϵ(a+b\epsilon)(c+d\epsilon)=ac+(ad+bc)\epsilon. The coefficient of ϵ\epsilon automatically follows the product rule.

Starting with x+1ϵx+1\epsilon, evaluate a program using these rules. The result is f(x)+f′(x)ϵf(x)+f'(x)\epsilon. In several dimensions, seed a direction vv to compute JvJv, a Jacobian-vector product. One pass computes one input direction; obtaining all columns generally takes nn seeds.

Reverse mode: carry an adjoint

Take a=x2a=x^2, b=3xb=3x, L=a+bL=a+b. A forward pass records x,a,b,Lx,a,b,L. Start backward with Lˉ=1\bar L=1. Addition sends 1 to both aa and bb. Their contributions to xx are 2x2x and 3, so xˉ=2x+3\bar x=2x+3.

A shared variable must accumulate contributions. Replacing one with another loses a path. Reverse mode computes JTvJ^Tv for an output seed vv. For one scalar loss, a single seed gives all parameter derivatives, at a cost comparable to a small number of forward evaluations. Storing or recomputing intermediate values is the main memory tradeoff.

Neither mode chooses a finite-difference step. It differentiates the operations on the executed path, subject to floating-point arithmetic and the derivative conventions of those operations. Discrete decisions and nondifferentiable points still need attention.

Quick check +20 XP

A node feeds two later operations. How does reverse mode combine their contributions?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Calculus, Volume 3

OpenStax · Section 4.5: the chain rule

Track each path through a dependency diagram.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 5.3–5.6: vector derivatives and backpropagation

Read Jacobian dimensions before multiplying matrices.

Book · free online · ~20 min

The Matrix Calculus You Need For Deep Learning

Parr & Howard · Jacobians and vector chain rules

Expand a small Jacobian one entry at a time.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Automatic Differentiation in Machine Learning: a SurveyBaydin, Pearlmutter, Radul & Siskind · 2018

The survey distinguishes automatic differentiation from symbolic algebra and finite differences. Reverse accumulation reuses intermediate values and sums contributions from all consumers of a node. The displayed rule restates that process in graph notation.

Decode the paper · Reverse accumulation, written for a general computation graph

Automatic Differentiation in Machine Learning: a Survey

Baydin, Pearlmutter, Radul & Siskind · 2018

+25 XP
vˉi=∑j: vi→vjvˉj∂vj∂vi\bar v_i=\sum_{j:\,v_i\to v_j}\bar v_j\frac{\partial v_j}{\partial v_i}

The survey distinguishes automatic differentiation from symbolic algebra and finite differences. Reverse accumulation reuses intermediate values and sums contributions from all consumers of a node. The displayed rule restates that process in graph notation.

vˉi\bar v_i
vi→vjv_i\to v_j
∂vj/∂vi\partial v_j/\partial v_i

Options

The Jacobian matrixKhan Academy
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Compare a nonlinear image of a small square with the parallelogram predicted by its Jacobian. Shrink the square, then locate three points where the pleat map has zero determinant.

Interactive lab

Zoom until it’s linear

The true image of a small square approaches the parallelogram predicted by the Jacobian. The plot automatically zooms around the image.
F(x,y)=(x, y3+xy)F(x, y) = (x,\ y^3 + xy)
True and linearised images of a small input square0.0793-0.2960.540.16510.6251.461.091.921.55First outputSecond output
● True boundary● Jacobian prediction
J=[1.0000.0000.5001.750]J=\begin{bmatrix}1.000&0.000\\0.500&1.750\end{bmatrix}

det J

1.7500

Boundary error / square side

0.345

Use the Pleat map for the challenge. Points must have |det J| ≤ 0.05 and be at least one unit apart.

Challenge: Fold finderFind three well-separated points where the map folds space flat, where det J = 0.+40 XP

Match · Expression ↔ Meaning

Jacobian shapes

+20 XP
F:R3→R2F:\mathbb R^3\to\mathbb R^2
F:R4→RF:\mathbb R^4\to\mathbb R
F:R→R5F:\mathbb R\to\mathbb R^5

Options

Match · Expression ↔ Meaning

Differentiation methods

+20 XP
JvJv
JTvJ^Tv
[f(x+h)−f(x−h)]/(2h)[f(x+h)-f(x-h)]/(2h)

Options

Proof puzzle

Compose local maps

+25 XP

Claim

Derive JG∘F=JGJFJ_{G\circ F}=J_GJ_F.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Dual multiplication is differentiation

+35 XP

Claim

Show that the epsilon coefficient of (f+f′ϵ)(g+g′ϵ)(f+f'\epsilon)(g+g'\epsilon) is the derivative of fgfg.

Preview

Your typeset proof appears here.

Coding problems

Problem 13·Warm-up

A composed Jacobian

+20 XP

Let F(x,y)=(x+y,xy)F(x,y)=(x+y,xy) and G(u,v)=(u2,v+u)G(u,v)=(u^2,v+u). Find the determinant of JG∘FJ_{G\circ F} at (2,3)(2,3).

An exact integer (or a fraction like 7/12)

Problem 14·Standard

Dual-number powers

+35 XP

Propagate the dual number 1+ϵ1+\epsilon through f(x)=∑k=1100xkf(x)=\sum_{k=1}^{100}x^k. Report the coefficient of epsilon.

An exact integer (or a fraction like 7/12)

Problem 15·Challenge

Count fold points

+50 XP

For F(x,y)=(x,y3+xy)F(x,y)=(x,y^3+xy), count integer pairs with −300≤x≤0-300\le x\le0 and −10≤y≤10-10\le y\le10 for which det⁡JF=0\det J_F=0.

An exact integer (or a fraction like 7/12)

Key takeaways

  • The Jacobian has outputs as rows and inputs as columns.
  • Forward mode computes Jv; reverse mode computes the transposed product.
  • Shared nodes accumulate derivative contributions.
  • A zero determinant concerns the local linear map and needs careful interpretation.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

For F=(x2+y,xy)F=(x^2+y,xy), what is det⁡JF\det J_F at (1,2)(1,2)?

Question 2 of 6 +20 XP

Which mode usually suits a scalar loss and a million inputs?

Question 3 of 6 +20 XP

What is the epsilon coefficient of (2+3ϵ)(4+5ϵ)(2+3\epsilon)(4+5\epsilon)?

Question 4 of 6 +20 XP

For L=x2+3xL=x^2+3x at x=2x=2, what adjoint reaches x?

Question 5 of 6 +20 XP

Does zero Jacobian determinant prove a map is globally non-invertible?

Question 6 of 6 +20 XP

What does forward mode with seed v compute?

End of the chamber

Clear this chamber

+60 XPJacobianMultivariable Chain RuleDual NumbersReverse Mode