Skip to content
AriadneTechnology

The Outer Ring · Chamber 3 of 9

The Chain Rule: Deep Compositions

Why derivatives multiply along a chain, proved properly, and why deep chains make gradients vanish or explode.

40 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Prove the chain rule and apply it to deep compositions
  • Differentiate inverse functions and logarithms
  • Use the log-derivative trick behind policy gradients
  • Explain vanishing and exploding gradients, and why residual connections help
DiscoverLearnRead beyondPapers & lecturesYour turn

A deep network is a composition. Each layer changes how much an earlier input perturbation matters. Residual networks add a path that passes the input directly to the output of a block; differentiating that path explains part of their behaviour.

Spotted in the wild

y=F(x,{Wi})+x\mathbf{y}=\mathcal F(\mathbf{x},\{W_i\})+\mathbf{x}
Deep Residual Learning for Image Recognition
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • f∘gf\circ g“f composed with g”
    Apply gg first, then ff.
  • f′(g(x))f'(g(x))“f prime evaluated at g of x”
    Outer derivative at the inner output.
  • dhLdh0\frac{dh_L}{dh_0}“d h L by d h zero”
    Sensitivity through the full chain.
  • ∏kak\prod_k a_k“product of a k”
    Multiply the local sensitivities.
  • f−1f^{-1}“inverse function”
    The function that undoes ff, where it exists.
  • ∇θlog⁡pθ\nabla_\theta\log p_\theta“gradient of log p”
    The score of a differentiable probability model.

Multiply local sensitivities

If y=f(u)y=f(u) and u=g(x)u=g(x), then

dydx=f′(g(x))g′(x).\frac{dy}{dx}=f'(g(x))g'(x).

For y=(3x+1)2y=(3x+1)^2, the outer slope is 2(3x+1)2(3x+1) and the inner slope is 3, giving 6(3x+1)6(3x+1). Evaluate each derivative at the input its function actually receives.

Quick check +20 XP

For y=(3x+1)2y=(3x+1)^2, what is y′(1)y'(1)?

A proof that handles zero slopes

Differentiability of ff at u=g(a)u=g(a) means

f(u+k)=f(u)+f′(u)k+r(k)k,r(k)→0.f(u+k)=f(u)+f'(u)k+r(k)k,\qquad r(k)\to0.

Set k=g(a+h)−g(a)k=g(a+h)-g(a). Differentiability of gg makes k→0k\to0 and k/h→g′(a)k/h\to g'(a). Divide by hh to get a limit of [f′(u)+r(k)]k/h=f′(u)g′(a)[f'(u)+r(k)]k/h=f'(u)g'(a). When k=0k=0, define the remainder contribution as zero. This avoids dividing by kk, which could be zero even for nonzero hh.

Inverses and logarithms

If ff has a differentiable local inverse and f′(x)≠0f'(x)\ne0, differentiating f−1(f(x))=xf^{-1}(f(x))=x gives (f−1)′(f(x))=1/f′(x)(f^{-1})'(f(x))=1/f'(x). Thus the derivative of ln⁡x\ln x, the inverse of exe^x, is 1/x1/x for x>0x>0.

For a positive differentiable function pθp_\theta, the log-derivative identity is

∇θpθ=pθ∇θlog⁡pθ.\nabla_\theta p_\theta=p_\theta\nabla_\theta\log p_\theta.

For a finite fixed outcome space and rewards R(x)R(x) independent of θ\theta, this yields

∇θ∑xpθ(x)R(x)=Epθ[R(X)∇θlog⁡pθ(X)].\nabla_\theta\sum_xp_\theta(x)R(x)=\mathbb E_{p_\theta}[R(X)\nabla_\theta\log p_\theta(X)].

This is the score-function route to policy gradients. Continuous or parameter-dependent supports require care about differentiating under the integral and boundary terms.

Quick check +20 XP

If p=0.2p=0.2 and dp/dθ=0.06dp/d\theta=0.06, what is dlog⁡p/dθd\log p/d\theta?

Twenty multiplications later

For hk=σ(wkhk−1+bk)h_k=\sigma(w_kh_{k-1}+b_k),

dhLdh0=∏k=1Lwkσ′(wkhk−1+bk).\frac{dh_L}{dh_0}=\prod_{k=1}^L w_k\sigma'(w_kh_{k-1}+b_k).

Twenty factors of 1/21/2 give about one millionth; twenty factors of 2 give over a million. Sigmoid's derivative is at most 1/41/4, but the weights also matter. It is incorrect to conclude that every sigmoid network's gradient must shrink solely from its activation bound.

A residual scalar layer hk+1=hk+Fk(hk)h_{k+1}=h_k+F_k(h_k) contributes 1+Fk′(hk)1+F_k'(h_k). Small residual derivatives can keep products near one. If Fk′=−1F_k'=-1, however, the total is zero. Residual connections help preserve paths for information; they do not abolish the chain rule.

For branched computations, multiply along each path and add across paths. In y=x2+xy=x^2+x, the two contributions are 2x2x and 1.

Quick check +20 XP

What happens to derivatives when two paths meet in a sum?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Calculus, Volume 1

OpenStax · Section 3.6: the chain rule

Identify the inner and outer function in each worked example.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 5.1–5.2: chain rules

Follow the move from a scalar chain to a network of dependencies.

Book · free online · ~20 min

The Matrix Calculus You Need For Deep Learning

Parr & Howard · The chain rule

Compare a straight chain with a graph that branches.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Deep Residual Learning for Image RecognitionKaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2016

The residual block adds an identity shortcut to a learned transformation. In the scalar analogue, its derivative is 1+F′(x)1+F'(x). This gives a direct gradient contribution, but does not guarantee a nonzero total: contributions can cancel, and later activations still matter.

Decode the paper · Equation (1): residual block

Deep Residual Learning for Image Recognition

Kaiming He, Xiangyu Zhang, Shaoqing Ren & Jian Sun · 2016

+25 XP
y=F(x,{Wi})+x\mathbf{y}=\mathcal F(\mathbf{x},\{W_i\})+\mathbf{x}

The residual block adds an identity shortcut to a learned transformation. In the scalar analogue, its derivative is 1+F′(x)1+F'(x). This gives a direct gradient contribution, but does not guarantee a nonzero total: contributions can cancel, and later activations still matter.

x\mathbf{x}
F\mathcal F
{Wi}\{W_i\}
y\mathbf{y}

Options

Visualizing the chain rule and product rule3Blue1Brown
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Build twenty identical scalar layers. Track both their activations and the product of their local derivatives, then tune the gain to preserve a useful gradient.

Interactive lab

Keep the gradient alive

Twenty layers use hₖ₊₁ = activation(gain × hₖ) − activation(0). Centring keeps zero a fixed point. Each derivative includes the gain.
Log magnitude of the derivative through twenty layers0-165-11.510-715-2.5202Layerlog₁₀ |derivative| (clipped at −16)
● Accumulated derivative● Lower target● Upper target

Final output

4.552e-14

Full derivative

9.089e-13

Challenge: Keep the gradient aliveKeep the end-to-end derivative of a 20-layer chain between 0.1 and 10 for three activations.+40 XP

Match · Expression ↔ Meaning

Choose the rule

+20 XP
f(g(x))f(g(x))
f(x)+g(x)f(x)+g(x)
f−1(f(x))f^{-1}(f(x))

Options

Match · Expression ↔ Meaning

Gradient products

+20 XP
0.5200.5^{20}
2202^{20}
(1+0)20(1+0)^{20}

Options

Proof puzzle

Differentiate an inverse

+25 XP

Claim

Show (f−1)′(f(x))=1/f′(x)(f^{-1})'(f(x))=1/f'(x) when the derivatives exist and f′(x)≠0f'(x)\ne0.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

A bound on deep derivatives

+35 XP

Claim

If each of LL differentiable scalar layers has derivative magnitude at most cc, prove that the composition has derivative magnitude at most cLc^L.

Preview

Your typeset proof appears here.

Coding problems

Problem 7·Warm-up

A shrinking chain

+20 XP

Every layer has derivative 0.8. Find the smallest number of layers for which the product is below 0.01.

An exact integer (or a fraction like 7/12)

Problem 8·Standard

Differentiate a recurrence

+35 XP

Start with h0=xh_0=x, then hk+1=hk2+1h_{k+1}=h_k^2+1 for three layers. Compute dh3/dxdh_3/dx at x=1x=1.

An exact integer (or a fraction like 7/12)

Problem 9·Challenge

A residual telescope

+50 XP

Layer k=1,…,100k=1,\ldots,100 is hk=hk−1+hk−1/kh_k=h_{k-1}+h_{k-1}/k. What is dh100/dh0dh_{100}/dh_0?

An exact integer (or a fraction like 7/12)

Key takeaways

  • Evaluate each local derivative at the right intermediate value.
  • Multiply along paths and add contributions where paths meet.
  • Inverse and log derivatives follow from the same chain rule.
  • Residual paths help gradient flow, but cancellation and large products remain possible.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

For y=e2xy=e^{2x}, what is y′(0)y'(0)?

Question 2 of 6 +20 XP

Five local derivatives are each 2. What is the full derivative?

Question 3 of 6 +20 XP

A residual branch has derivative −0.2-0.2. What is the derivative of x+F(x)x+F(x)?

Question 4 of 6 +20 XP

When can the derivative of an inverse be computed as a reciprocal?

Question 5 of 6 +20 XP

Why can a deep composition have vanishing gradients?

Question 6 of 6 +20 XP

For y=x2+xy=x^2+x at x=2x=2, what is dy/dxdy/dx?

End of the chamber

Clear this chamber

+60 XPCompositionInverse DerivativeLog-Derivative IdentityGradient Products