A residual layer adds a small change to its current state. Let those changes become a continuous velocity field and the network becomes a differential equation. Recovering its final state requires accumulating change over time.
Spotted in the wild
- “delta x”Width of one integration strip.
- “integral of f from a to b”Signed accumulation over an interval.
- “F prime equals f”An antiderivative relation.
- “absolute Jacobian determinant”Local volume factor in a change of variables.
- “order n to the minus two”Error bounded by a constant times n to the minus two for large n.
- “z dot”Time derivative of the state.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “delta x” | Width of one integration strip. | ||
| “integral of f from a to b” | Signed accumulation over an interval. | ||
| “F prime equals f” | An antiderivative relation. | ||
| “absolute Jacobian determinant” | Local volume factor in a change of variables. | ||
| “order n to the minus two” | Error bounded by a constant times n to the minus two for large n. | ||
| “z dot” | Time derivative of the state. |
An area is a limit of sums
Partition into equal intervals of width . Pick within each strip. For a continuous function,
The integral is signed: contributions below the axis are negative. To measure unsigned area, integrate . For on , the right sum is , tending to .
What is ?
The fundamental theorem
Define for continuous f. Then
The average value over the shrinking interval tends to by continuity. Thus . Conversely, if , the definite integral is .
This converts accumulation into differentiation in reverse. It also differentiates integrals with moving boundaries: by the chain rule.
What is the derivative of at x=2?
Substitution and integration by parts
For a differentiable change , substitution gives . Definite integrals need transformed bounds. The formula comes directly from the chain rule.
Integrating the product rule gives . For example, .
In multiple dimensions, volume changes by the absolute Jacobian determinant. Polar coordinates have area element . This factor is essential in the Gaussian integral. Let . Then, using positivity to combine integrals,
Since , . Rescaling gives the standard normal normalising constant .
Numerical accumulation
For sufficiently smooth functions and a fixed interval, left and right rules generally have error ; midpoint and trapezoid have ; composite Simpson has and requires an even number of strips. Endpoint singularities can spoil these rates.
A Monte Carlo estimator uses uniform on : . It is unbiased when the expectation exists, and has standard error proportional to when the variance is finite. Its strength is handling high dimensions, rather than fast convergence for smooth one-dimensional curves.
From residual steps to an ODE
Euler's method for is . A residual update has this form. For with , n equal steps to time 1 give , approaching . Finer discretisation improves this example, but some ODEs require much smaller steps for stability.
For z prime = z, z(0)=1, one Euler step of size 0.1 gives which state?
Read beyond
Book · free online · ~20 min
Calculus, Volume 1OpenStax · Sections 5.2–5.5: definite integrals and substitution
Read both parts of the fundamental theorem and their assumptions.
Book · free online · ~20 min
Calculus, Volume 2OpenStax · Sections 3.1 and 3.6: integration by parts and numerical integration
Compare an exact antiderivative with a numerical estimate.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Section 6.2: probability distributions
Interpret integrals as continuous weighted sums.
Read the equation in context
Neural Ordinary Differential EquationsRicky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt & David Duvenaud · 2018Neural ODEs define the hidden-state evolution through a learned derivative and use an ODE solver to obtain later states. The integral contains the unknown trajectory z(t); it is not generally an ordinary integral of a known fixed curve. Solver tolerance and computational cost become part of evaluating the model.
Decode the paper · Equation (2): continuous hidden-state dynamics
Neural Ordinary Differential EquationsRicky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt & David Duvenaud · 2018
Neural ODEs define the hidden-state evolution through a learned derivative and use an ODE solver to obtain later states. The integral contains the unknown trajectory z(t); it is not generally an ordinary integral of a known fixed curve. Solver tolerance and computational cost become part of evaluating the model.
Options
Your turn
Compare left, midpoint, trapezoid and Simpson estimates. Reach an absolute error below 0.001 for three integrals with at most sixteen strips each.
Interactive lab
Riemann racer
Estimate
1.512437
Exact integral
1.718282
Absolute error
2.058e-1
Match · Expression ↔ Meaning
Integration rules
Options
Match · Expression ↔ Meaning
From calculus to models
Options
Proof puzzle
Accumulation has a derivative
Claim
For continuous f, show has derivative f(x).
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Unbiased Monte Carlo integration
Claim
For independent uniform U_i on [a,b] and integrable f, prove is unbiased for the integral.
Your typeset proof appears here.
Coding problems
Problem 25·Warm-up
A right Riemann sum
Approximate with 100 right-endpoint strips. Submit the result as a reduced fraction.
Problem 26·Standard
Euler’s network
Start z=1 and repeat one hundred times. Report the final state to 6 decimal places.
Problem 27·Challenge
Find an integration budget
Use composite midpoint integration for on . Find the smallest positive n for which the absolute error is strictly below .
Key takeaways
- Integrals accumulate signed change through limits of sums.
- The fundamental theorem connects accumulation to derivatives.
- Substitution and multidimensional changes of variables require the correct scale factor.
- Residual steps can approximate continuous dynamics, with numerical error and stability to monitor.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
An integral over a curve below the horizontal axis contributes what?
What is ?
Which factor appears in polar area integration?
Composite Simpson’s rule requires which strip count?
How many times more independent samples reduce Monte Carlo standard error by half?
Why is a neural ODE integral more involved than integrating a known curve?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Riemann racer (+40 XP)
- Bonus: Problem 25: A right Riemann sum (+20 XP)
- Bonus: Problem 26: Euler’s network (+35 XP)
- Bonus: Problem 27: Find an integration budget (+50 XP)
- Bonus: Proof: Accumulation has a derivative (+25 XP)
- Bonus: Proof: Unbiased Monte Carlo integration (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Integration rules (+20 XP)
- Bonus: Match: From calculus to models (+20 XP)