A loss surface may depend on millions of parameters. A two-dimensional slice lets us inspect a few directions through that space. The gradient tells us how the loss changes along any chosen direction at the point where we stand.
Spotted in the wild
- “partial f by partial x i”Rate of change with other coordinates fixed.
- “gradient of f”Column of partial derivatives.
- “directional derivative along u”Slope per unit distance when is a unit vector.
- “u has Euclidean norm one”A fair constraint when comparing directions.
- “gradient dot u”First-order change along the direction.
- “level set of f at c”Points having the same function value.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “partial f by partial x i” | Rate of change with other coordinates fixed. | ||
| “gradient of f” | Column of partial derivatives. | ||
| “directional derivative along u” | Slope per unit distance when is a unit vector. | ||
| “u has Euclidean norm one” | A fair constraint when comparing directions. | ||
| “gradient dot u” | First-order change along the direction. | ||
| “level set of f at c” | Points having the same function value. |
Hold everything else fixed
For , a partial derivative changes one coordinate while holding the other fixed:
At the gradient is . A small displacement changes the function by approximately . We use column gradients throughout this course.
For , what is at ?
Move in any direction
For a unit vector , follow the path . The chain rule gives the directional derivative
For and gradient , the slope is . A direction must be normalised when comparing slopes per unit distance; doubling an unnormalised vector doubles the parameterised rate of change.
Partial derivatives alone do not always imply differentiability. The dot-product formula assumes a genuine local linear approximation. Continuous partial derivatives in a neighbourhood are a useful sufficient condition.
If and , what is ?
Why the gradient points uphill
Cauchy–Schwarz gives . When the gradient is nonzero, equality occurs at . Negate it for steepest descent.
This statement uses Euclidean distance. Rescaling parameters changes what counts as a unit step and can change the optimisation path. Preconditioners exploit this by changing how directions are scaled.
At a stationary point, , every first-order directional derivative is zero. Curvature, which we study later, is needed to distinguish a minimum from a saddle.
Gradients meet contours at right angles
Suppose a differentiable path stays on a level set . Differentiating gives . Thus a nonzero gradient is perpendicular to the tangent of a smooth contour.
For , the gradient points radially outward. A tangent has zero dot product with it.
Taking a finite step
Gradient descent uses . Its first-order change is , negative at a nonstationary point. The neglected error means the step must still be small enough.
For , the update is . It converges to zero for , oscillates without shrinking at , and diverges for larger positive steps. A direction can be downhill locally while a long step raises the loss.
Why can a gradient step raise the loss?
Read beyond
Book · free online · ~20 min
Calculus, Volume 3OpenStax · Sections 4.3–4.6: partial and directional derivatives
Compute a gradient and compare it with tangent directions to a contour.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Section 5.2: partial differentiation and gradients
Check the dimensions of each gradient.
Book · free online · ~20 min
Convex OptimizationBoyd & Vandenberghe · Section 9.3: gradient descent
Read the descent argument and the role of the step size.
Read the equation in context
Visualizing the Loss Landscape of Neural NetsHao Li, Zheng Xu, Gavin Taylor, Christoph Studer & Tom Goldstein · 2018The paper plots losses along selected directions and studies how normalising those directions changes the view. A flat-looking slice need not imply flatness in every direction. The plotted surface depends on the chosen plane and the scale of its axes.
Decode the paper · Section 3: two-dimensional visualisation
Visualizing the Loss Landscape of Neural NetsHao Li, Zheng Xu, Gavin Taylor, Christoph Studer & Tom Goldstein · 2018
The paper plots losses along selected directions and studies how normalising those directions changes the view. A flat-looking slice need not imply flatness in every direction. The plotted surface depends on the chosen plane and the scale of its axes.
Options
Your turn
Choose a point, then aim your direction arrow down the loss surface. Compare the directional derivative with the best possible descent slope, and find a direction tangent to a contour.
Interactive lab
Downhill compass
Directional derivative
-15.910
Best unit descent slope
-16.862
Descent: 0/3 · Tangent: remaining
Match · Expression ↔ Meaning
Slopes and vectors
Options
Match · Expression ↔ Meaning
At a contour
Options
Proof puzzle
Steepest ascent
Claim
For nonzero , prove the largest slope along a unit direction is .
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
A gradient is normal to a level set
Claim
If is constant and are differentiable, prove .
Your typeset proof appears here.
Coding problems
Problem 10·Warm-up
A directional derivative
For , evaluate the directional derivative at along the unit vector . Give one decimal place.
Problem 11·Standard
Walk down the bowl
Run gradient descent on , starting at with step size 0.1. After 20 updates, report to 8 decimal places.
Problem 12·Challenge
Count downhill grid directions
At a point with gradient , consider all integer vectors with , excluding . How many give a strictly negative directional derivative after normalisation?
Key takeaways
- Gradients collect partial derivatives; directional derivatives are dot products.
- The gradient gives steepest ascent under a Euclidean unit-step constraint.
- Contour tangents are orthogonal to a nonzero gradient.
- A downhill direction still needs an appropriate step size.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
If , what is the largest unit-direction slope?
If and , what is ?
At a differentiable stationary point, which is guaranteed?
For , and , what is the next gradient descent iterate?
What makes the gradient perpendicular to a contour?
What does a two-dimensional loss slice show?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Downhill compass (+30 XP)
- Bonus: Problem 10: A directional derivative (+20 XP)
- Bonus: Problem 11: Walk down the bowl (+35 XP)
- Bonus: Problem 12: Count downhill grid directions (+50 XP)
- Bonus: Proof: Steepest ascent (+25 XP)
- Bonus: Proof: A gradient is normal to a level set (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Slopes and vectors (+20 XP)
- Bonus: Match: At a contour (+20 XP)