A network maps a vector of parameters to a scalar loss. Writing out every derivative would be wasteful. Automatic differentiation applies local derivative rules to the executed computation and propagates only the products needed for that loss.
Spotted in the wild
- “Jacobian of F”Output-by-input matrix of partial derivatives.
- “J times v”Forward directional derivative of a vector output.
- “J transpose times v”Backward propagation of an output sensitivity.
- “v bar”Adjoint: derivative of the scalar loss with respect to v.
- “epsilon squared equals zero”The defining dual-number rule; epsilon is formal, not a tiny float.
- “determinant of J”Signed local volume scaling for a square Jacobian.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “Jacobian of F” | Output-by-input matrix of partial derivatives. | ||
| “J times v” | Forward directional derivative of a vector output. | ||
| “J transpose times v” | Backward propagation of an output sensitivity. | ||
| “v bar” | Adjoint: derivative of the scalar loss with respect to v. | ||
| “epsilon squared equals zero” | The defining dual-number rule; epsilon is formal, not a tiny float. | ||
| “determinant of J” | Signed local volume scaling for a square Jacobian. |
One derivative for every input-output pair
For , the Jacobian has shape :
Rows correspond to outputs and columns to inputs. If , then
At , a perturbation predicts output changes . The first output's exact change has an additional from the square.
For , how many entries are in ?
Composition becomes multiplication
For and ,
The shapes are . The order follows the forward composition. For a scalar loss after , column gradients travel backward as .
A square Jacobian's determinant measures signed local area or volume scaling. A zero determinant means the linear approximation loses a dimension. It does not by itself prove global non-invertibility: is one-to-one even though its derivative vanishes at zero.
If is a composition, which is its Jacobian?
Forward mode: carry a tangent
A dual number has . Multiplication gives . The coefficient of automatically follows the product rule.
Starting with , evaluate a program using these rules. The result is . In several dimensions, seed a direction to compute , a Jacobian-vector product. One pass computes one input direction; obtaining all columns generally takes seeds.
Reverse mode: carry an adjoint
Take , , . A forward pass records . Start backward with . Addition sends 1 to both and . Their contributions to are and 3, so .
A shared variable must accumulate contributions. Replacing one with another loses a path. Reverse mode computes for an output seed . For one scalar loss, a single seed gives all parameter derivatives, at a cost comparable to a small number of forward evaluations. Storing or recomputing intermediate values is the main memory tradeoff.
Neither mode chooses a finite-difference step. It differentiates the operations on the executed path, subject to floating-point arithmetic and the derivative conventions of those operations. Discrete decisions and nondifferentiable points still need attention.
A node feeds two later operations. How does reverse mode combine their contributions?
Read beyond
Book · free online · ~20 min
Calculus, Volume 3OpenStax · Section 4.5: the chain rule
Track each path through a dependency diagram.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 5.3–5.6: vector derivatives and backpropagation
Read Jacobian dimensions before multiplying matrices.
Book · free online · ~20 min
The Matrix Calculus You Need For Deep LearningParr & Howard · Jacobians and vector chain rules
Expand a small Jacobian one entry at a time.
Read the equation in context
Automatic Differentiation in Machine Learning: a SurveyBaydin, Pearlmutter, Radul & Siskind · 2018The survey distinguishes automatic differentiation from symbolic algebra and finite differences. Reverse accumulation reuses intermediate values and sums contributions from all consumers of a node. The displayed rule restates that process in graph notation.
Decode the paper · Reverse accumulation, written for a general computation graph
Automatic Differentiation in Machine Learning: a SurveyBaydin, Pearlmutter, Radul & Siskind · 2018
The survey distinguishes automatic differentiation from symbolic algebra and finite differences. Reverse accumulation reuses intermediate values and sums contributions from all consumers of a node. The displayed rule restates that process in graph notation.
Options
Your turn
Compare a nonlinear image of a small square with the parallelogram predicted by its Jacobian. Shrink the square, then locate three points where the pleat map has zero determinant.
Interactive lab
Zoom until it’s linear
det J
1.7500
Boundary error / square side
0.345
Use the Pleat map for the challenge. Points must have |det J| ≤ 0.05 and be at least one unit apart.
Match · Expression ↔ Meaning
Jacobian shapes
Options
Match · Expression ↔ Meaning
Differentiation methods
Options
Proof puzzle
Compose local maps
Claim
Derive .
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Dual multiplication is differentiation
Claim
Show that the epsilon coefficient of is the derivative of .
Your typeset proof appears here.
Coding problems
Problem 13·Warm-up
A composed Jacobian
Let and . Find the determinant of at .
Problem 14·Standard
Dual-number powers
Propagate the dual number through . Report the coefficient of epsilon.
Problem 15·Challenge
Count fold points
For , count integer pairs with and for which .
Key takeaways
- The Jacobian has outputs as rows and inputs as columns.
- Forward mode computes Jv; reverse mode computes the transposed product.
- Shared nodes accumulate derivative contributions.
- A zero determinant concerns the local linear map and needs careful interpretation.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
For , what is at ?
Which mode usually suits a scalar loss and a million inputs?
What is the epsilon coefficient of ?
For at , what adjoint reaches x?
Does zero Jacobian determinant prove a map is globally non-invertible?
What does forward mode with seed v compute?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Fold finder (+40 XP)
- Bonus: Problem 13: A composed Jacobian (+20 XP)
- Bonus: Problem 14: Dual-number powers (+35 XP)
- Bonus: Problem 15: Count fold points (+50 XP)
- Bonus: Proof: Compose local maps (+25 XP)
- Bonus: Proof: Dual multiplication is differentiation (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Jacobian shapes (+20 XP)
- Bonus: Match: Differentiation methods (+20 XP)