A computer usually supplies uniform pseudorandom numbers. Sampling a waiting time, a Gaussian or a categorical choice requires transforming that starting noise. The Gumbel-max trick turns independent continuous noise into a discrete draw.
Spotted in the wild
- “mass at k”Probability of a discrete outcome.
- “density at x”Probability per unit width for a continuous variable.
- “CDF at x”Probability of being at most x.
- “quantile at level u”The smallest point with cumulative probability at least u.
- “X is normally distributed”A Gaussian with specified mean and variance.
- “n choose k”Number of k-element subsets of n trials.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “mass at k” | Probability of a discrete outcome. | ||
| “density at x” | Probability per unit width for a continuous variable. | ||
| “CDF at x” | Probability of being at most x. | ||
| “quantile at level u” | The smallest point with cumulative probability at least u. | ||
| “X is normally distributed” | A Gaussian with specified mean and variance. | ||
| “n choose k” | Number of k-element subsets of n trials. |
Mass, density and cumulative probability
A discrete probability mass function gives , with . A continuous density instead gives interval probabilities by integration: . At a single point the probability is zero, and a density may exceed one.
The cumulative distribution function works in both settings. It is nondecreasing and right-continuous, tending to zero and one at the two extremes. A discrete CDF jumps by the probability mass; a continuous differentiable CDF has derivative p.
A density at a point is 2.5. Is that possible?
Distributions to recognise
| Model | Support and probability rule | Mean and variance |
|---|---|---|
| Bernoulli(p) | 0 or 1; success chance p | p, p(1-p) |
| Binomial(n,p) | k successes; | np, np(1-p) |
| Geometric(p) | Trials to first success, k=1,2,…; | 1/p, (1-p)/p² |
| Poisson(λ) | k=0,1,…; | λ, λ |
| Exponential(λ) | x≥0; | 1/λ, 1/λ² |
| Normal(μ,σ²) | Real x; | μ, σ² |
Binomial probabilities count which k trials succeed, then multiply the probability of each ordered outcome. A geometric waiting time requires k−1 failures before success. Taking the binomial limit with n large and np fixed yields the Poisson model. For a constant event rate λ, the chance of no event by time x is , giving the exponential CDF .
Always state whether a geometric variable counts trials or failures. The two conventions differ by one.
Invert the CDF
For a continuous strictly increasing CDF and , define . Then
For distributions with jumps use the generalised inverse . This produces discrete categories too.
An exponential draw is . The equally distributed is valid as well. Avoid endpoints that make a logarithm undefined in numerical code. A categorical draw selects the first cumulative probability at least U.
An exponential has rate 2. What quantile corresponds to ?
Transform a density
For a differentiable one-to-one transformation with nonzero derivative on its domain,
Probability is preserved while interval widths change. For X uniform on [0,1] and Y=X², the CDF is on [0,1], so the density is in the interior. It exceeds one near zero but still integrates to one.
A many-to-one transformation requires adding contributions from all inverse branches. Squaring a variable that can be positive or negative is one such case. In several dimensions the inverse derivative becomes an absolute Jacobian determinant.
If X is uniform on [0,1] and Y=X², what is ?
Read beyond
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapters 3–4: discrete and continuous random variables
Compare the roles of a mass function, density and CDF.
Book · free online · ~20 min
Introduction to ProbabilityBlitzstein & Hwang · Random variables, named distributions and transformations
Keep track of the support of each named distribution.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Sections 6.1–6.2 and 6.7: distributions and transformations
Check the Jacobian factor in a transformed density.
Read the equation in context
Categorical Reparameterization with Gumbel-SoftmaxEric Jang, Shixiang Gu & Ben Poole · 2017The paper begins with categorical sampling by adding Gumbel noise to log probabilities, then replaces argmax with a temperature-controlled softmax. The hard draw is discrete; the relaxed vector is differentiable but is not an exact one-hot draw at positive temperature.
Decode the paper · Section 2: Gumbel-max sampling, before the continuous relaxation
Categorical Reparameterization with Gumbel-SoftmaxEric Jang, Shixiang Gu & Ben Poole · 2017
The paper begins with categorical sampling by adding Gumbel noise to log probabilities, then replaces argmax with a temperature-controlled softmax. The hard draw is discrete; the relaxed vector is differentiable but is not an exact one-hot draw at positive temperature.
Options
Your turn
Use the CDF to map a probability level to a quantile. Change the distribution and observe how uniformly spaced probability levels produce unevenly spaced samples.
Interactive lab
Inverse transform sampler
Gaps between requests to a server are Exp(λ = 2) seconds. Find the median gap: half of all gaps are shorter than it.
Quantile F⁻¹(u)
0.14384
Match · Expression ↔ Meaning
Distribution roles
Options
Match · Expression ↔ Meaning
Sampling recipes
Options
Proof puzzle
Inverse-transform sampling
Claim
For continuous strictly increasing F, show F inverse of U has CDF F.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
A transformed uniform variable
Claim
Derive the CDF and density of Y=X² for X uniform on [0,1].
Your typeset proof appears here.
Coding problems
Problem 4·Warm-up
Exactly three successes
For X binomial with n=8 and p=1/2, compute P(X=3) as a reduced fraction.
Problem 5·Standard
An exponential quantile
For an exponential distribution with rate 2, compute the 95th percentile to 6 decimal places.
Problem 6·Challenge
A rare binomial tail
For X binomial with n=20 and p=0.1, compute P(X≥5) to 8 decimal places.
Key takeaways
- A density can exceed one; a probability cannot.
- State distribution support and parameter conventions.
- CDF inversion transforms uniform samples into a target distribution.
- Density transformations preserve probability by accounting for local volume.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
A binomial variable has n=10 and p=0.3. What is its mean?
A geometric variable counts trials through the first success, with p=0.25. What is its mean?
What happens to a CDF at a discrete atom?
What is the variance of a Poisson variable with rate 4?
A change of variables has two inverse branches. What must a density calculation do?
What is the exact Gumbel-max output?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Quantile quest (+30 XP)
- Bonus: Problem 4: Exactly three successes (+20 XP)
- Bonus: Problem 5: An exponential quantile (+35 XP)
- Bonus: Problem 6: A rare binomial tail (+50 XP)
- Bonus: Proof: Inverse-transform sampling (+25 XP)
- Bonus: Proof: A transformed uniform variable (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Distribution roles (+20 XP)
- Bonus: Match: Sampling recipes (+20 XP)