Skip to content
AriadneTechnology

The Outer Ring · Chamber 1 of 9

Limits: Getting Arbitrarily Close

What it means to approach without arriving: ε–δ, continuity, and the convergent series behind SGD's step sizes.

40 min 50 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • State and use the ε–δ definition of a limit
  • Decide where a function is continuous, and see why a step can't be trained by gradients
  • Sum geometric series and tell convergent series from divergent ones
  • Read the Robbins–Monro conditions on stochastic gradient descent's step sizes
DiscoverLearnRead beyondPapers & lecturesYour turn

A learning rate can shrink to zero while its total travel remains unbounded. That distinction lets a noisy optimiser keep correcting its estimate without giving every new sample the same influence. To understand it, we need two different limits: the limit of a sequence and the limit of its partial sums.

Spotted in the wild

∑n=1∞an=∞,∑n=1∞an2<∞\sum_{n=1}^{\infty}a_n=\infty,\qquad \sum_{n=1}^{\infty}a_n^2<\infty
A Stochastic Approximation Method
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • lim⁡x→af(x)\lim_{x\to a}f(x)“limit of f as x approaches a”
    The value approached near aa.
  • ε\varepsilon“epsilon”
    Requested output accuracy.
  • δ\delta“delta”
    An input radius chosen for that accuracy.
  • ∣x−a∣|x-a|“distance from x to a”
    Absolute error in the input.
  • SNS_N“S N”
    The sum of the first NN terms.
  • ∑n=1∞an\sum_{n=1}^{\infty}a_n“sum of a n from one to infinity”
    The limit of partial sums, if it exists.

A limit describes nearby values

The statement lim⁡x→af(x)=L\lim_{x\to a}f(x)=L concerns inputs near aa, with x≠ax\ne a. The value at aa may be different or undefined. For example,

x2−1x−1=x+1(x≠1),lim⁡x→1x2−1x−1=2.\frac{x^2-1}{x-1}=x+1\quad(x\ne1),\qquad \lim_{x\to1}\frac{x^2-1}{x-1}=2.

Cancelling is valid away from the missing point, which is exactly where the limit looks. A two-sided limit exists only when the left and right limits agree.

Quick check +20 XP

What is lim⁡x→1(x2−1)/(x−1)\lim_{x\to1}(x^2-1)/(x-1)?

Make “close” precise

For every desired output error ε>0\varepsilon>0, there must be an input radius δ>0\delta>0 such that

0<∣x−a∣<δ⟹∣f(x)−L∣<ε.0<|x-a|<\delta\quad\Longrightarrow\quad |f(x)-L|<\varepsilon.

The order matters: the challenger chooses ε\varepsilon, then you choose δ\delta. Your choice must work for all eligible xx, not just the points drawn on a screen.

For f(x)=3x+1f(x)=3x+1 at a=2a=2, the target is L=7L=7. Since ∣f(x)−7∣=3∣x−2∣|f(x)-7|=3|x-2|, choose δ=ε/3\delta=\varepsilon/3.

For f(x)=x2f(x)=x^2 at a=2a=2, factor the error: ∣x2−4∣=∣x−2∣∣x+2∣|x^2-4|=|x-2||x+2|. First require δ≤1\delta\le1, so ∣x+2∣<5|x+2|<5. Then δ=min⁡(1,ε/5)\delta=\min(1,\varepsilon/5) works. A valid radius need not be the largest one.

Quick check +20 XP

For f(x)=3x+1f(x)=3x+1, choose δ=ε/3\delta=\varepsilon/3. What is δ\delta when ε=0.06\varepsilon=0.06?

Continuity and a broken learning signal

A function is continuous at aa if its limit there equals f(a)f(a). Differentiability will be a stronger condition: ∣x∣|x| is continuous at zero but has a corner.

A hard threshold s(x)=0s(x)=0 for x<0x<0 and s(x)=1s(x)=1 for x≥0x\ge0 jumps at zero. It has derivative zero away from zero and no derivative at zero. Ordinary gradient descent therefore gets no useful local signal through this activation. Smooth or piecewise linear activations give us more useful slopes.

A vanishing term is not a convergent sum

A sequence ana_n has a limit. A series ∑an\sum a_n means the limit of partial sums SN=∑n=1NanS_N=\sum_{n=1}^N a_n. Those are different questions.

For ∣r∣<1|r|<1, subtract rSNrS_N from SNS_N to obtain

∑k=0N−1rk=1−rN1−r,∑k=0∞rk=11−r.\sum_{k=0}^{N-1}r^k=\frac{1-r^N}{1-r},\qquad \sum_{k=0}^{\infty}r^k=\frac1{1-r}.

The remaining tail after NN terms is rN/(1−r)r^N/(1-r) for 0<r<10<r<1. This gives an actual stopping rule for code.

The harmonic series diverges even though 1/n→01/n\to0: group terms from 2k−1+12^{k-1}+1 to 2k2^k. Each group contributes at least 1/21/2, forever. By comparison with the integral of x−px^{-p}, the positive series ∑n−p\sum n^{-p} converges exactly when p>1p>1.

Consequently an=n−pa_n=n^{-p} satisfies both step-size conditions when 1/2<p≤11/2<p\le1. Squaring doubles the exponent. A geometric schedule has a finite total sum, so it fails the first condition even though it shrinks smoothly.

Quick check +20 XP

Which schedule satisfies both ∑an=∞\sum a_n=\infty and ∑an2<∞\sum a_n^2<\infty?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Calculus, Volume 1

OpenStax · Sections 2.2–2.5: limits and continuity

Work through a two-sided limit before reading the formal definition.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Section 5.1: differentiation of univariate functions

Trace how a limit turns a finite slope into a derivative.

Book · free online · ~20 min

Convex Optimization

Boyd & Vandenberghe · Chapter 9: unconstrained minimisation

Compare shrinking steps with line search, which chooses steps using the objective.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

A Stochastic Approximation MethodHerbert Robbins & Sutton Monro · 1951

Robbins and Monro study finding a root from noisy observations. These conditions balance continued movement against accumulated noise. They are part of a convergence theorem with assumptions on the response and its noise; choosing these steps alone does not guarantee that an arbitrary neural network converges.

Decode the paper · Step-size assumptions, with the paper’s notation a_n

A Stochastic Approximation Method

Herbert Robbins & Sutton Monro · 1951

+25 XP
∑n=1∞an=∞,∑n=1∞an2<∞\sum_{n=1}^{\infty}a_n=\infty,\qquad \sum_{n=1}^{\infty}a_n^2<\infty

Robbins and Monro study finding a root from noisy observations. These conditions balance continued movement against accumulated noise. They are part of a convergence theorem with assumptions on the response and its noise; choosing these steps alone does not guarantee that an arbitrary neural network converges.

ana_n
∑an\sum a_n
∑an2\sum a_n^2

Options

Limits, L'Hôpital's rule, and epsilon delta definitions3Blue1Brown
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Choose an input radius that guarantees the requested output accuracy. A graph can suggest a choice; an inequality certifies every point in the interval.

Interactive lab

The ε–δ game

For f(x)=x² at x=2, keep every value with |x−2|<δ inside |f(x)−4|<ε.

Round 1 of 3 · ε = 0.5

A quadratic near x=2 with upper and lower output tolerances1.722.911.863.46242.144.542.285.09xy
● x²● 4 + ε● 4 − ε

Supremum of the output error

0.8379

Challenge: The ε–δ gameAnswer every ε with a δ that keeps the curve inside the band, and spot the limit that doesn't exist.+40 XP

Match · Expression ↔ Meaning

Limits and sums

+20 XP
an→0a_n\to0
SN→SS_N\to S
f(a)=lim⁡x→af(x)f(a)=\lim_{x\to a}f(x)

Options

Match · Expression ↔ Meaning

Which series converges?

+20 XP
∑1/n\sum 1/n
∑1/n2\sum 1/n^2
∑n=0∞(1/2)n\sum_{n=0}^{\infty}(1/2)^n

Options

Proof puzzle

Sum a geometric series

+25 XP

Claim

For ∣r∣<1|r|<1, prove ∑k=0∞rk=1/(1−r)\sum_{k=0}^{\infty}r^k=1/(1-r).

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Certify a quadratic limit

+35 XP

Claim

Prove lim⁡x→2x2=4\lim_{x\to2}x^2=4 using the epsilon–delta definition.

Preview

Your typeset proof appears here.

Coding problems

Problem 1·Warm-up

How many geometric steps?

+20 XP

Let SN=∑k=0N−1(1/2)kS_N=\sum_{k=0}^{N-1}(1/2)^k. Find the smallest positive integer NN for which 2−SN<10−62-S_N<10^{-6}.

An exact integer (or a fraction like 7/12)

Problem 2·Standard

Telescoping travel

+35 XP

Compute ∑n=1100001/[n(n+1)]\sum_{n=1}^{10000}1/[n(n+1)]. Submit a fraction in lowest terms.

An exact integer (or a fraction like 7/12)

Problem 3·Challenge

Audit the schedules

+50 XP

For each integer k=1,…,200k=1,\ldots,200, consider an=n−k/100a_n=n^{-k/100}. How many schedules satisfy both Robbins–Monro sum conditions?

An exact integer (or a fraction like 7/12)

Key takeaways

  • Limits concern nearby values; continuity also checks the value at the point.
  • An epsilon–delta proof must work for every input in its interval.
  • Terms tending to zero are necessary, but insufficient, for a convergent series.
  • Power schedules satisfy both sum conditions exactly for 1/2<p≤11/2<p\le1.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

A function has a limit at a missing point. What follows?

Question 2 of 6 +20 XP

Compute ∑k=0∞(1/4)k\sum_{k=0}^{\infty}(1/4)^k.

Question 3 of 6 +20 XP

Why does 1/n→01/n\to0 fail to prove ∑1/n\sum 1/n converges?

Question 4 of 6 +20 XP

Which is the correct order in the limit definition?

Question 5 of 6 +20 XP

What is lim⁡x→2x2\lim_{x\to2}x^2?

Question 6 of 6 +20 XP

Why does a hard threshold hinder ordinary gradient learning?

End of the chamber

Clear this chamber

+50 XPLimitEpsilon and DeltaContinuityGeometric Series