Two weights with the same magnitude can have very different effects when removed. A quadratic approximation of the loss estimates the damage using curvature. That idea connects Taylor expansions with neural-network pruning.
Spotted in the wild
- “Taylor remainder”Error after a degree-n Taylor polynomial.
- “Hessian”Matrix of second partial derivatives.
- “h transpose H h”Curvature contribution along a displacement.
- “eigenvalue i of H”Curvature along a Hessian eigenvector.
- “H h equals minus g”Linear system for a Newton displacement.
- “H is positive definite”Every nonzero direction has positive quadratic curvature.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “Taylor remainder” | Error after a degree-n Taylor polynomial. | ||
| “Hessian” | Matrix of second partial derivatives. | ||
| “h transpose H h” | Curvature contribution along a displacement. | ||
| “eigenvalue i of H” | Curvature along a Hessian eigenvector. | ||
| “H h equals minus g” | Linear system for a Newton displacement. | ||
| “H is positive definite” | Every nonzero direction has positive quadratic curvature. |
Approximate more than the slope
A Taylor polynomial matches derivatives at a point:
If throughout the interval and the usual Taylor theorem hypotheses hold, then . This gives an error guarantee for a finite approximation. Smoothness alone does not guarantee that an infinite Taylor series equals the function.
For around zero, the first three terms are . At , they give 1.105. A cubic remainder bound is .
What is the degree-two Taylor polynomial of at zero evaluated at ?
The Hessian measures directional curvature
For a scalar function with continuous second partial derivatives, the Hessian is symmetric:
For , . The mixed partial records how the slope in one coordinate changes as another moves. Along a unit direction , the second derivative is .
Classify a stationary point
At , Hessian eigenvalues reveal the quadratic behaviour:
| Eigenvalues | Conclusion |
|---|---|
| All strictly positive | Strict local minimum |
| All strictly negative | Strict local maximum |
| At least one positive and one negative | Saddle point |
| Some zero, with no mixed signs | Test inconclusive |
The zero Hessian at zero cannot distinguish from . A positive semidefinite Hessian at one point is insufficient to prove a local minimum.
At a stationary point, Hessian eigenvalues are 2 and -1. What is it?
Newton follows the quadratic model
Set the gradient of the quadratic approximation to zero: . Solve this linear system for and update . Writing describes the formula; implementations usually solve the system without forming an inverse.
For , Newton reaches 3 from any initial point in one step. For a nonquadratic function, full steps can overshoot. An indefinite Hessian can even send Newton toward a maximum or saddle. Damping, line search, or replacing by a positive definite approximation can help.
A badly conditioned positive Hessian makes gradient descent zigzag: one step size must accommodate both steep and shallow directions. Newton rescales by curvature, but computing and storing a dense Hessian is expensive in a large model.
Predict the cost of deleting a weight
Deleting means setting and other displacements to zero. The Taylor prediction is . Near a stationary point the first term is small, leaving the saliency in the opening equation. A small weight on a very curved direction can matter more than a larger weight on a flat direction.
If , and , what is the deletion saliency?
Read beyond
Book · free online · ~20 min
Calculus, Volume 2OpenStax · Chapter 6: power series
Distinguish a finite Taylor polynomial from an infinite series.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 5.7–5.8: higher derivatives and Taylor series
Write the quadratic approximation with its dimensions visible.
Book · free online · ~20 min
Convex OptimizationBoyd & Vandenberghe · Section 9.5: Newton’s method
Read how damping and line search improve the full Newton step.
Read the equation in context
Optimal Brain DamageYann LeCun, John S. Denker & Sara A. Solla · 1989Optimal Brain Damage estimates deletion cost using second derivatives. The displayed saliency assumes training is near a stationary point and uses a diagonal quadratic approximation. Correlations between simultaneous deletions and higher-order effects can make the approximation inaccurate.
Decode the paper · Diagonal second-order saliency, with H used for the loss Hessian
Optimal Brain DamageYann LeCun, John S. Denker & Sara A. Solla · 1989
Optimal Brain Damage estimates deletion cost using second derivatives. The displayed saliency assumes training is near a stationary point and uses a diagonal quadratic approximation. Correlations between simultaneous deletions and higher-order effects can make the approximation inaccurate.
Options
Your turn
Explore Himmelblau’s function. Locate stationary points and classify them from the two Hessian eigenvalues. Positive, negative and mixed curvature give different local geometry.
Interactive lab
Critical point cartographer
Gradient norm
5.97e+1
Hessian eigenvalues
-29.31, -6.69
Match · Expression ↔ Meaning
Stationary geometry
Options
Match · Expression ↔ Meaning
Approximation terms
Options
Proof puzzle
Newton’s displacement
Claim
Derive the Newton step from for symmetric invertible H.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
A quadratic remainder
Claim
For , prove the linear approximation error at a is exactly .
Your typeset proof appears here.
Coding problems
Problem 19·Warm-up
Approximate the exponential
Compute , the degree-ten Taylor approximation to , to 8 decimal places.
Problem 20·Standard
Newton’s reciprocal
Use Newton’s method to minimise for , starting at . After four full updates report x to 8 decimal places.
Problem 21·Challenge
Rank deletion costs
For weights indexed by , let and . Sum the predicted saliencies of deleting weights 51 through 100, treating each deletion separately. Give 6 decimal places.
Key takeaways
- A finite Taylor polynomial needs a remainder estimate to certify accuracy.
- Hessian eigenvalues classify nondegenerate stationary points.
- Zero eigenvalues require further analysis.
- Newton solves a local quadratic model; curvature estimates also suggest pruning costs.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
For , starting at x=0, where does a full Newton step go?
A stationary point has a zero Hessian. What follows?
For , what is ?
What is preferable to explicitly forming an inverse for a Newton step?
For and , what is ?
When does a finite Taylor error bound apply?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Critical cartographer (+50 XP)
- Bonus: Problem 19: Approximate the exponential (+20 XP)
- Bonus: Problem 20: Newton’s reciprocal (+35 XP)
- Bonus: Problem 21: Rank deletion costs (+50 XP)
- Bonus: Proof: Newton’s displacement (+25 XP)
- Bonus: Proof: A quadratic remainder (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Stationary geometry (+20 XP)
- Bonus: Match: Approximation terms (+20 XP)