Skip to content
AriadneTechnology

The Outer Ring · Chamber 2 of 9

Distributions and How to Sample Them

PMFs, PDFs and CDFs, the named distributions of machine learning, and how to turn uniform noise into any distribution you like.

40 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Move between PMFs, PDFs and CDFs, and read off quantiles
  • Derive and use the binomial, geometric, Poisson, exponential and normal distributions
  • Sample any distribution by inverse transform, including the Gumbel-max trick
  • Transform a density through a change of variables
DiscoverLearnRead beyondPapers & lecturesYour turn

A computer usually supplies uniform pseudorandom numbers. Sampling a waiting time, a Gaussian or a categorical choice requires transforming that starting noise. The Gumbel-max trick turns independent continuous noise into a discrete draw.

Spotted in the wild

k=argmax⁡i(log⁡πi+gi),gi=−log⁡(−log⁡Ui)k=\operatorname*{argmax}_i(\log\pi_i+g_i),\qquad g_i=-\log(-\log U_i)
Categorical Reparameterization with Gumbel-Softmax
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • p(k)p(k)“mass at k”
    Probability of a discrete outcome.
  • p(x)p(x)“density at x”
    Probability per unit width for a continuous variable.
  • F(x)F(x)“CDF at x”
    Probability of being at most x.
  • F−1(u)F^{-1}(u)“quantile at level u”
    The smallest point with cumulative probability at least u.
  • X∼N(μ,σ2)X\sim\mathcal N(\mu,\sigma^2)“X is normally distributed”
    A Gaussian with specified mean and variance.
  • (nk)\binom nk“n choose k”
    Number of k-element subsets of n trials.

Mass, density and cumulative probability

A discrete probability mass function gives p(k)=P(X=k)p(k)=P(X=k), with ∑kp(k)=1\sum_kp(k)=1. A continuous density instead gives interval probabilities by integration: P(a<X≤b)=∫abp(x)dxP(a<X\le b)=\int_a^bp(x)dx. At a single point the probability is zero, and a density may exceed one.

The cumulative distribution function F(x)=P(X≤x)F(x)=P(X\le x) works in both settings. It is nondecreasing and right-continuous, tending to zero and one at the two extremes. A discrete CDF jumps by the probability mass; a continuous differentiable CDF has derivative p.

Quick check +20 XP

A density at a point is 2.5. Is that possible?

Distributions to recognise

ModelSupport and probability ruleMean and variance
Bernoulli(p)0 or 1; success chance pp, p(1-p)
Binomial(n,p)k successes; (nk)pk(1−p)n−k\binom nkp^k(1-p)^{n-k}np, np(1-p)
Geometric(p)Trials to first success, k=1,2,…; (1−p)k−1p(1-p)^{k-1}p1/p, (1-p)/p²
Poisson(λ)k=0,1,…; e−λλk/k!e^{-\lambda}\lambda^k/k!λ, λ
Exponential(λ)x≥0; λe−λx\lambda e^{-\lambda x}1/λ, 1/λ²
Normal(μ,σ²)Real x; e−(x−μ)2/(2σ2)/(σ2π)e^{-(x-\mu)^2/(2\sigma^2)}/(\sigma\sqrt{2\pi})μ, σ²

Binomial probabilities count which k trials succeed, then multiply the probability of each ordered outcome. A geometric waiting time requires k−1 failures before success. Taking the binomial limit with n large and np fixed yields the Poisson model. For a constant event rate λ, the chance of no event by time x is e−λxe^{-\lambda x}, giving the exponential CDF 1−e−λx1-e^{-\lambda x}.

Always state whether a geometric variable counts trials or failures. The two conventions differ by one.

Invert the CDF

For a continuous strictly increasing CDF and U∼Uniform⁡(0,1)U\sim\operatorname{Uniform}(0,1), define X=F−1(U)X=F^{-1}(U). Then

P(X≤x)=P(U≤F(x))=F(x).P(X\le x)=P(U\le F(x))=F(x).

For distributions with jumps use the generalised inverse F−1(u)=inf⁡{x:F(x)≥u}F^{-1}(u)=\inf\{x:F(x)\ge u\}. This produces discrete categories too.

An exponential draw is X=−log⁡(1−U)/λX=-\log(1-U)/\lambda. The equally distributed −log⁡U/λ-\log U/\lambda is valid as well. Avoid endpoints that make a logarithm undefined in numerical code. A categorical draw selects the first cumulative probability at least U.

Quick check +20 XP

An exponential has rate 2. What quantile corresponds to u=1−e−2u=1-e^{-2}?

Transform a density

For a differentiable one-to-one transformation Y=g(X)Y=g(X) with nonzero derivative on its domain,

pY(y)=pX(g−1(y))∣ddyg−1(y)∣.p_Y(y)=p_X(g^{-1}(y))\left|\frac{d}{dy}g^{-1}(y)\right|.

Probability is preserved while interval widths change. For X uniform on [0,1] and Y=X², the CDF is y\sqrt y on [0,1], so the density is 1/(2y)1/(2\sqrt y) in the interior. It exceeds one near zero but still integrates to one.

A many-to-one transformation requires adding contributions from all inverse branches. Squaring a variable that can be positive or negative is one such case. In several dimensions the inverse derivative becomes an absolute Jacobian determinant.

Quick check +20 XP

If X is uniform on [0,1] and Y=X², what is P(Y≤1/4)P(Y\le1/4)?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapters 3–4: discrete and continuous random variables

Compare the roles of a mass function, density and CDF.

Book · free online · ~20 min

Introduction to Probability

Blitzstein & Hwang · Random variables, named distributions and transformations

Keep track of the support of each named distribution.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 6.1–6.2 and 6.7: distributions and transformations

Check the Jacobian factor in a transformed density.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Categorical Reparameterization with Gumbel-SoftmaxEric Jang, Shixiang Gu & Ben Poole · 2017

The paper begins with categorical sampling by adding Gumbel noise to log probabilities, then replaces argmax with a temperature-controlled softmax. The hard draw is discrete; the relaxed vector is differentiable but is not an exact one-hot draw at positive temperature.

Decode the paper · Section 2: Gumbel-max sampling, before the continuous relaxation

Categorical Reparameterization with Gumbel-Softmax

Eric Jang, Shixiang Gu & Ben Poole · 2017

+25 XP
k=argmax⁡i(log⁡πi+gi),gi=−log⁡(−log⁡Ui)k=\operatorname*{argmax}_i(\log\pi_i+g_i),\qquad g_i=-\log(-\log U_i)

The paper begins with categorical sampling by adding Gumbel noise to log probabilities, then replaces argmax with a temperature-controlled softmax. The hard draw is discrete; the relaxed vector is differentiable but is not an exact one-hot draw at positive temperature.

πi\pi_i
UiU_i
gig_i
kk

Options

The Main Ideas behind Probability DistributionsStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Use the CDF to map a probability level to a quantile. Change the distribution and observe how uniformly spaced probability levels produce unevenly spaced samples.

Interactive lab

Inverse transform sampler

Choose a probability level and follow it across to the CDF, then down to a quantile. Uniform probability draws become samples by the same rule.

Gaps between requests to a server are Exp(λ = 2) seconds. Find the median gap: half of all gaps are shorter than it.

Cumulative distribution with a selected probability and quantile000.9250.251.850.52.780.753.71xP(X ≤ x)
● CDF● Probability level● Quantile intersection

Quantile F⁻¹(u)

0.14384

Challenge: Quantile questRead three quantiles straight off the CDF, each to within 0.01.+30 XP

Match · Expression ↔ Meaning

Distribution roles

+20 XP
p(k)p(k) for discrete X
∫abp(x)dx\int_a^bp(x)dx
F−1(u)F^{-1}(u)

Options

Match · Expression ↔ Meaning

Sampling recipes

+20 XP
−log⁡U/λ-\log U/\lambda
μ+σZ\mu+\sigma Z, Z standard normal
arg⁡max⁡i(log⁡πi+gi)\arg\max_i(\log\pi_i+g_i)

Options

Proof puzzle

Inverse-transform sampling

+25 XP

Claim

For continuous strictly increasing F, show F inverse of U has CDF F.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

A transformed uniform variable

+35 XP

Claim

Derive the CDF and density of Y=X² for X uniform on [0,1].

Preview

Your typeset proof appears here.

Coding problems

Problem 4·Warm-up

Exactly three successes

+20 XP

For X binomial with n=8 and p=1/2, compute P(X=3) as a reduced fraction.

An exact integer (or a fraction like 7/12)

Problem 5·Standard

An exponential quantile

+35 XP

For an exponential distribution with rate 2, compute the 95th percentile to 6 decimal places.

A number, rounded to 6 decimal places

Problem 6·Challenge

A rare binomial tail

+50 XP

For X binomial with n=20 and p=0.1, compute P(X≥5) to 8 decimal places.

A number, rounded to 8 decimal places

Key takeaways

  • A density can exceed one; a probability cannot.
  • State distribution support and parameter conventions.
  • CDF inversion transforms uniform samples into a target distribution.
  • Density transformations preserve probability by accounting for local volume.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

A binomial variable has n=10 and p=0.3. What is its mean?

Question 2 of 6 +20 XP

A geometric variable counts trials through the first success, with p=0.25. What is its mean?

Question 3 of 6 +20 XP

What happens to a CDF at a discrete atom?

Question 4 of 6 +20 XP

What is the variance of a Poisson variable with rate 4?

Question 5 of 6 +20 XP

A change of variables has two inverse branches. What must a density calculation do?

Question 6 of 6 +20 XP

What is the exact Gumbel-max output?

End of the chamber

Clear this chamber

+60 XPCumulative DistributionBinomial DistributionInverse Transform SamplingDensity Transformation