Skip to content
AriadneTechnology

The Middle Ring · Chamber 4 of 9

Gradients: Steepest Ascent in Many Dimensions

Partial derivatives, directional derivatives and contour maps, and a proof that the gradient points straight uphill.

40 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Compute partial derivatives and gradients of functions of many variables
  • Compute directional derivatives and prove the gradient is the direction of steepest ascent
  • Show that the gradient is perpendicular to level sets
  • Read the loss-landscape plots of research papers
DiscoverLearnRead beyondPapers & lecturesYour turn

A loss surface may depend on millions of parameters. A two-dimensional slice lets us inspect a few directions through that space. The gradient tells us how the loss changes along any chosen direction at the point where we stand.

Spotted in the wild

f(α,β)=L(θ∗+αδ+βη)f(\alpha,\beta)=L(\theta^*+\alpha\delta+\beta\eta)
Visualizing the Loss Landscape of Neural Nets
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • ∂f/∂xi\partial f/\partial x_i“partial f by partial x i”
    Rate of change with other coordinates fixed.
  • ∇f\nabla f“gradient of f”
    Column of partial derivatives.
  • DufD_uf“directional derivative along u”
    Slope per unit distance when uu is a unit vector.
  • ∥u∥2=1\|u\|_2=1“u has Euclidean norm one”
    A fair constraint when comparing directions.
  • ∇fTu\nabla f^Tu“gradient dot u”
    First-order change along the direction.
  • {x:f(x)=c}\{x:f(x)=c\}“level set of f at c”
    Points having the same function value.

Hold everything else fixed

For f(x,y)=x2+3y2f(x,y)=x^2+3y^2, a partial derivative changes one coordinate while holding the other fixed:

∂f∂x=2x,∂f∂y=6y,∇f=(2x,6y)T.\frac{\partial f}{\partial x}=2x,\qquad \frac{\partial f}{\partial y}=6y,\qquad \nabla f=(2x,6y)^T.

At (1,1)(1,1) the gradient is (2,6)T(2,6)^T. A small displacement h=(h1,h2)Th=(h_1,h_2)^T changes the function by approximately 2h1+6h22h_1+6h_2. We use column gradients throughout this course.

Quick check +20 XP

For f=x2+3y2f=x^2+3y^2, what is ∂f/∂y\partial f/\partial y at (1,1)(1,1)?

Move in any direction

For a unit vector uu, follow the path x(t)=x+tux(t)=x+tu. The chain rule gives the directional derivative

Duf(x)=ddtf(x+tu)∣t=0=∇f(x)Tu.D_uf(x)=\left.\frac{d}{dt}f(x+tu)\right|_{t=0}=\nabla f(x)^Tu.

For u=(3/5,4/5)u=(3/5,4/5) and gradient (2,6)(2,6), the slope is 6/5+24/5=66/5+24/5=6. A direction must be normalised when comparing slopes per unit distance; doubling an unnormalised vector doubles the parameterised rate of change.

Partial derivatives alone do not always imply differentiability. The dot-product formula assumes a genuine local linear approximation. Continuous partial derivatives in a neighbourhood are a useful sufficient condition.

Quick check +20 XP

If ∇f=(2,6)\nabla f=(2,6) and u=(3/5,4/5)u=(3/5,4/5), what is DufD_uf?

Why the gradient points uphill

Cauchy–Schwarz gives ∇fTu≤∥∇f∥∥u∥=∥∇f∥\nabla f^Tu\le\|\nabla f\|\|u\|=\|\nabla f\|. When the gradient is nonzero, equality occurs at u=∇f/∥∇f∥u=\nabla f/\|\nabla f\|. Negate it for steepest descent.

This statement uses Euclidean distance. Rescaling parameters changes what counts as a unit step and can change the optimisation path. Preconditioners exploit this by changing how directions are scaled.

At a stationary point, ∇f=0\nabla f=0, every first-order directional derivative is zero. Curvature, which we study later, is needed to distinguish a minimum from a saddle.

Gradients meet contours at right angles

Suppose a differentiable path c(t)c(t) stays on a level set f(c(t))=kf(c(t))=k. Differentiating gives ∇f(c(t))Tc′(t)=0\nabla f(c(t))^Tc'(t)=0. Thus a nonzero gradient is perpendicular to the tangent of a smooth contour.

For x2+y2=kx^2+y^2=k, the gradient (2x,2y)(2x,2y) points radially outward. A tangent (−y,x)(-y,x) has zero dot product with it.

Taking a finite step

Gradient descent uses xt+1=xt−η∇f(xt)x_{t+1}=x_t-\eta\nabla f(x_t). Its first-order change is −η∥∇f∥2-\eta\|\nabla f\|^2, negative at a nonstationary point. The neglected error means the step must still be small enough.

For f(x)=x2f(x)=x^2, the update is xt+1=(1−2η)xtx_{t+1}=(1-2\eta)x_t. It converges to zero for 0<η<10<\eta<1, oscillates without shrinking at η=1\eta=1, and diverges for larger positive steps. A direction can be downhill locally while a long step raises the loss.

Quick check +20 XP

Why can a gradient step raise the loss?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Calculus, Volume 3

OpenStax · Sections 4.3–4.6: partial and directional derivatives

Compute a gradient and compare it with tangent directions to a contour.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Section 5.2: partial differentiation and gradients

Check the dimensions of each gradient.

Book · free online · ~20 min

Convex Optimization

Boyd & Vandenberghe · Section 9.3: gradient descent

Read the descent argument and the role of the step size.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Visualizing the Loss Landscape of Neural NetsHao Li, Zheng Xu, Gavin Taylor, Christoph Studer & Tom Goldstein · 2018

The paper plots losses along selected directions and studies how normalising those directions changes the view. A flat-looking slice need not imply flatness in every direction. The plotted surface depends on the chosen plane and the scale of its axes.

Decode the paper · Section 3: two-dimensional visualisation

Visualizing the Loss Landscape of Neural Nets

Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer & Tom Goldstein · 2018

+25 XP
f(α,β)=L(θ∗+αδ+βη)f(\alpha,\beta)=L(\theta^*+\alpha\delta+\beta\eta)

The paper plots losses along selected directions and studies how normalising those directions changes the view. A flat-looking slice need not imply flatness in every direction. The plotted surface depends on the chosen plane and the scale of its axes.

θ∗\theta^*
δ,η\delta,\eta
α,β\alpha,\beta
LL

Options

Gradient Descent, Step-by-StepStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Choose a point, then aim your direction arrow down the loss surface. Compare the directional derivative with the best possible descent slope, and find a direction tangent to a contour.

Interactive lab

Downhill compass

Aim within 3° of steepest descent at three distinct points, then find a contour-tangent direction at another point.
L(w,b)=14∑i=14(wxi+b−yi)2L(w, b) = \tfrac{1}{4}\sum_{i=1}^{4} (wx_i + b - y_i)^2
Loss contours, a chosen direction and the gradient-1.7-1.2-0.450.050.81.32.052.553.33.8wb
● Contours● Your direction● Gradient direction● Current point

Directional derivative

-15.910

Best unit descent slope

-16.862

Descent: 0/3 · Tangent: remaining

Challenge: Downhill compassAim within 3° of steepest descent at three points, then find a direction with zero slope.+30 XP

Match · Expression ↔ Meaning

Slopes and vectors

+20 XP
∇f\nabla f
DufD_uf
−∇f-\nabla f

Options

Match · Expression ↔ Meaning

At a contour

+20 XP
∇fTu>0\nabla f^Tu>0
∇fTu<0\nabla f^Tu<0
∇fTu=0\nabla f^Tu=0

Options

Proof puzzle

Steepest ascent

+25 XP

Claim

For nonzero g=∇fg=\nabla f, prove the largest slope along a unit direction is ∥g∥\|g\|.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

A gradient is normal to a level set

+35 XP

Claim

If f(c(t))f(c(t)) is constant and f,cf,c are differentiable, prove ∇f(c(t))Tc′(t)=0\nabla f(c(t))^Tc'(t)=0.

Preview

Your typeset proof appears here.

Coding problems

Problem 10·Warm-up

A directional derivative

+20 XP

For f(x,y)=x2+xy+2y2f(x,y)=x^2+xy+2y^2, evaluate the directional derivative at (2,−1)(2,-1) along the unit vector (3/5,4/5)(3/5,4/5). Give one decimal place.

A number, rounded to 1 decimal place

Problem 11·Standard

Walk down the bowl

+35 XP

Run gradient descent on f(x,y)=x2+2y2f(x,y)=x^2+2y^2, starting at (1,1)(1,1) with step size 0.1. After 20 updates, report ff to 8 decimal places.

A number, rounded to 8 decimal places

Problem 12·Challenge

Count downhill grid directions

+50 XP

At a point with gradient (3,4)(3,4), consider all integer vectors (a,b)(a,b) with −20≤a,b≤20-20\le a,b\le20, excluding (0,0)(0,0). How many give a strictly negative directional derivative after normalisation?

An exact integer (or a fraction like 7/12)

Key takeaways

  • Gradients collect partial derivatives; directional derivatives are dot products.
  • The gradient gives steepest ascent under a Euclidean unit-step constraint.
  • Contour tangents are orthogonal to a nonzero gradient.
  • A downhill direction still needs an appropriate step size.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

If ∇f=(3,4)\nabla f=(3,4), what is the largest unit-direction slope?

Question 2 of 6 +20 XP

If ∇f=(3,4)\nabla f=(3,4) and u=(−4/5,3/5)u=(-4/5,3/5), what is DufD_uf?

Question 3 of 6 +20 XP

At a differentiable stationary point, which is guaranteed?

Question 4 of 6 +20 XP

For f=x2f=x^2, x=2x=2 and η=0.1\eta=0.1, what is the next gradient descent iterate?

Question 5 of 6 +20 XP

What makes the gradient perpendicular to a contour?

Question 6 of 6 +20 XP

What does a two-dimensional loss slice show?

End of the chamber

Clear this chamber

+60 XPPartial DerivativeDirectional DerivativeLevel SetSteepest Ascent