A deep network is a composition. Each layer changes how much an earlier input perturbation matters. Residual networks add a path that passes the input directly to the output of a block; differentiating that path explains part of their behaviour.
Spotted in the wild
- “f composed with g”Apply first, then .
- “f prime evaluated at g of x”Outer derivative at the inner output.
- “d h L by d h zero”Sensitivity through the full chain.
- “product of a k”Multiply the local sensitivities.
- “inverse function”The function that undoes , where it exists.
- “gradient of log p”The score of a differentiable probability model.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “f composed with g” | Apply first, then . | ||
| “f prime evaluated at g of x” | Outer derivative at the inner output. | ||
| “d h L by d h zero” | Sensitivity through the full chain. | ||
| “product of a k” | Multiply the local sensitivities. | ||
| “inverse function” | The function that undoes , where it exists. | ||
| “gradient of log p” | The score of a differentiable probability model. |
Multiply local sensitivities
If and , then
For , the outer slope is and the inner slope is 3, giving . Evaluate each derivative at the input its function actually receives.
For , what is ?
A proof that handles zero slopes
Differentiability of at means
Set . Differentiability of makes and . Divide by to get a limit of . When , define the remainder contribution as zero. This avoids dividing by , which could be zero even for nonzero .
Inverses and logarithms
If has a differentiable local inverse and , differentiating gives . Thus the derivative of , the inverse of , is for .
For a positive differentiable function , the log-derivative identity is
For a finite fixed outcome space and rewards independent of , this yields
This is the score-function route to policy gradients. Continuous or parameter-dependent supports require care about differentiating under the integral and boundary terms.
If and , what is ?
Twenty multiplications later
For ,
Twenty factors of give about one millionth; twenty factors of 2 give over a million. Sigmoid's derivative is at most , but the weights also matter. It is incorrect to conclude that every sigmoid network's gradient must shrink solely from its activation bound.
A residual scalar layer contributes . Small residual derivatives can keep products near one. If , however, the total is zero. Residual connections help preserve paths for information; they do not abolish the chain rule.
For branched computations, multiply along each path and add across paths. In , the two contributions are and 1.
What happens to derivatives when two paths meet in a sum?
Read beyond
Book · free online · ~20 min
Calculus, Volume 1OpenStax · Section 3.6: the chain rule
Identify the inner and outer function in each worked example.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 5.1–5.2: chain rules
Follow the move from a scalar chain to a network of dependencies.
Book · free online · ~20 min
The Matrix Calculus You Need For Deep LearningParr & Howard · The chain rule
Compare a straight chain with a graph that branches.
Read the equation in context
Deep Residual Learning for Image RecognitionKaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2016The residual block adds an identity shortcut to a learned transformation. In the scalar analogue, its derivative is . This gives a direct gradient contribution, but does not guarantee a nonzero total: contributions can cancel, and later activations still matter.
Decode the paper · Equation (1): residual block
Deep Residual Learning for Image RecognitionKaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2016
The residual block adds an identity shortcut to a learned transformation. In the scalar analogue, its derivative is . This gives a direct gradient contribution, but does not guarantee a nonzero total: contributions can cancel, and later activations still matter.
Options
Your turn
Build twenty identical scalar layers. Track both their activations and the product of their local derivatives, then tune the gain to preserve a useful gradient.
Interactive lab
Keep the gradient alive
Final output
4.552e-14
Full derivative
9.089e-13
Match · Expression ↔ Meaning
Choose the rule
Options
Match · Expression ↔ Meaning
Gradient products
Options
Proof puzzle
Differentiate an inverse
Claim
Show when the derivatives exist and .
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
A bound on deep derivatives
Claim
If each of differentiable scalar layers has derivative magnitude at most , prove that the composition has derivative magnitude at most .
Your typeset proof appears here.
Coding problems
Problem 7·Warm-up
A shrinking chain
Every layer has derivative 0.8. Find the smallest number of layers for which the product is below 0.01.
Problem 8·Standard
Differentiate a recurrence
Start with , then for three layers. Compute at .
Problem 9·Challenge
A residual telescope
Layer is . What is ?
Key takeaways
- Evaluate each local derivative at the right intermediate value.
- Multiply along paths and add contributions where paths meet.
- Inverse and log derivatives follow from the same chain rule.
- Residual paths help gradient flow, but cancellation and large products remain possible.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
For , what is ?
Five local derivatives are each 2. What is the full derivative?
A residual branch has derivative . What is the derivative of ?
When can the derivative of an inverse be computed as a reciprocal?
Why can a deep composition have vanishing gradients?
For at , what is ?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Keep the gradient alive (+40 XP)
- Bonus: Problem 7: A shrinking chain (+20 XP)
- Bonus: Problem 8: Differentiate a recurrence (+35 XP)
- Bonus: Problem 9: A residual telescope (+50 XP)
- Bonus: Proof: Differentiate an inverse (+25 XP)
- Bonus: Proof: A bound on deep derivatives (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Choose the rule (+20 XP)
- Bonus: Match: Gradient products (+20 XP)