Skip to content
AriadneTechnology

The Inner Ring · Chamber 8 of 9

Bayesian Inference: Priors, Posteriors and Bandits

Treat parameters as uncertain: conjugate priors, MAP estimates that turn out to be regularisation, and Thompson sampling.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Update a Beta prior with Bernoulli data, and a Gaussian prior with Gaussian data
  • Compare maximum-likelihood, MAP and posterior-mean estimates
  • Show that MAP estimation with a Gaussian prior is L2 regularisation
  • Balance exploration and exploitation with Thompson sampling
DiscoverLearnRead beyondPapers & lecturesYour turn

A bandit algorithm must decide whether to use the option that currently looks best or gather information about another. Thompson sampling makes that choice by sampling plausible reward probabilities from its current posterior beliefs.

Spotted in the wild

θa∼Beta⁡(αa,βa),A=argmax⁡aθa\theta_a\sim\operatorname{Beta}(\alpha_a,\beta_a),\qquad A=\operatorname*{argmax}_a\theta_a
A Tutorial on Thompson Sampling
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • p(θ)p(\theta)“parameter prior”
    Belief before observing the current data.
  • p(θ∣D)p(\theta\mid D)“parameter posterior”
    Updated distribution conditional on data.
  • Beta⁡(α,β)\operatorname{Beta}(\alpha,\beta)“Beta distribution”
    A density over Bernoulli success probabilities.
  • θ^MAP\hat\theta_{MAP}“maximum a posteriori estimate”
    A mode of the posterior density.
  • p(xnew∣D)p(x_{new}\mid D)“posterior predictive”
    Future-data distribution averaged over the posterior.
  • 1/σ21/\sigma^2“precision”
    Reciprocal variance for a scalar Gaussian.

Parameters can have distributions

Bayesian inference combines a prior p(θ)p(\theta) and likelihood p(D∣θ)p(D\mid\theta):

p(θ∣D)=p(D∣θ)p(θ)p(D).p(\theta\mid D)=\frac{p(D\mid\theta)p(\theta)}{p(D)}.

The posterior describes parameter uncertainty conditional on the model and data. A prior is part of that model. More data can reduce its influence, but only if the likelihood is informative about the parameter in question.

Beta plus Bernoulli

The Beta density on (0,1) has kernel θα−1(1−θ)β−1\theta^{\alpha-1}(1-\theta)^{\beta-1} for positive shape parameters. With s successes and f failures, multiply by the Bernoulli likelihood θs(1−θ)f\theta^s(1-\theta)^f:

θ∣D∼Beta⁡(α+s,β+f).\theta\mid D\sim\operatorname{Beta}(\alpha+s,\beta+f).

The posterior stays in the same family, which is called conjugacy. Its mean is (α+s)/(α+β+s+f)(\alpha+s)/(\alpha+\beta+s+f). For posterior shape parameters both greater than one, the mode is (α+s−1)/(α+β+s+f−2)(\alpha+s-1)/(\alpha+\beta+s+f-2). Boundary modes require separate treatment.

From a uniform Beta(1,1) prior, eight successes and two failures give Beta(9,3). The MLE is 0.8, the posterior mean is 0.75, and the interior MAP is 0.8. These answer different estimation questions.

Quick check +20 XP

A Beta(1,1) prior sees eight successes and two failures. What is the posterior mean?

Predict by averaging over uncertainty

The posterior predictive integrates parameters out:

p(xnew∣D)=∫p(xnew∣θ)p(θ∣D)dθ.p(x_{new}\mid D)=\int p(x_{new}\mid\theta)p(\theta\mid D)d\theta.

For one new Bernoulli trial, its success probability equals the posterior mean. For several future trials, using a fixed plug-in probability can miss dependence induced by the shared uncertain θ. The Beta-binomial predictive accounts for that uncertainty.

Gaussian conjugacy

Suppose μ∼N(μ0,τ2)\mu\sim\mathcal N(\mu_0,\tau^2) and observations are conditionally independent xi∣μ∼N(μ,σ2)x_i\mid\mu\sim\mathcal N(\mu,\sigma^2) with known σ². Completing the square gives posterior variance and mean

vn=(1τ2+nσ2)−1,mn=vn(μ0τ2+nxˉσ2).v_n=\left(\frac1{\tau^2}+\frac n{\sigma^2}\right)^{-1},\qquad m_n=v_n\left(\frac{\mu_0}{\tau^2}+\frac{n\bar x}{\sigma^2}\right).

Precision, the reciprocal variance, adds. The next observation has predictive variance σ2+vn\sigma^2+v_n, combining observation noise with remaining mean uncertainty.

Quick check +20 XP

A Gaussian posterior for a mean has variance 0.2. Observation noise variance is 1. What is the predictive variance of one new observation?

MAP and regularisation

The maximum a posteriori estimate maximises log likelihood plus log prior. An isotropic Gaussian prior on weights has negative log density equal to a constant plus ∥w∥2/(2τ2)\|w\|^2/(2\tau^2). Thus MAP minimises negative log-likelihood plus an L2 penalty. The penalty coefficient depends on whether the loss is summed or averaged.

MAP supplies one parameter value. The posterior mean and posterior predictive retain different aspects of uncertainty. A credible interval assigns posterior probability to a parameter region under the specified model; it is not automatically a frequentist coverage guarantee.

Thompson sampling

Maintain a Beta posterior for each machine, initially Beta(1,1). Draw one θ from each posterior, choose the largest, observe the reward, and increment the selected arm’s success or failure count. Broad posteriors sometimes produce optimistic draws, encouraging exploration.

No finite budget guarantees discovery of the best arm with a specified confidence for every possible reward configuration. Very similar machines can require many trials. A simulated posterior probability of being best is itself a numerical estimate, conditional on the model.

Quick check +20 XP

What makes Thompson sampling explore?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Sections 8.4 and 9.3: Bayesian inference

Compare a parameter point estimate with a posterior distribution.

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapter 9: Bayesian inference

Write the likelihood and prior kernels before identifying a conjugate posterior.

Book · free online · ~20 min

Introduction to Probability

Blitzstein & Hwang · Beta distributions and conjugacy

Connect Beta parameters with Bernoulli counts.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

A Tutorial on Thompson SamplingDaniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband & Zheng Wen · 2018

The tutorial presents Thompson sampling as choosing actions from sampled plausible models. In the Bernoulli example, successes and failures update Beta posteriors. The sampled winner is an exploration strategy, not proof that its true reward probability is highest.

Decode the paper · Beta-Bernoulli bandit example, written as a single action-selection step

A Tutorial on Thompson Sampling

Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband & Zheng Wen · 2018

+25 XP
θa∼Beta⁡(αa,βa),A=argmax⁡aθa\theta_a\sim\operatorname{Beta}(\alpha_a,\beta_a),\qquad A=\operatorname*{argmax}_a\theta_a

The tutorial presents Thompson sampling as choosing actions from sampled plausible models. In the Bernoulli example, successes and failures update Beta posteriors. The sampled winner is an exploration strategy, not proof that its true reward probability is highest.

aa
αa,βa\alpha_a,\beta_a
θa\theta_a
AA

Options

Bayes' Theorem, Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Pull three machines and update their Beta posteriors. Compare choosing the largest posterior mean with sampling a plausible success probability for each arm. Observe both uncertainty and accumulated reward.

Interactive lab

Three slot machines

Each machine has a hidden, fixed Bernoulli reward probability. Start with independent Beta(1,1) priors. Allocate up to sixty pulls manually or with Thompson sampling; a trial may not collect enough evidence to reach the target.

Machine A

0 rewards / 0 pulls
Posterior: Beta(1, 1)

Posterior mean

0.500

Probability of being best

33.33%

Machine B

0 rewards / 0 pulls
Posterior: Beta(1, 1)

Posterior mean

0.500

Probability of being best

33.33%

Machine C

0 rewards / 0 pulls
Posterior: Beta(1, 1)

Posterior mean

0.500

Probability of being best

33.33%

Pulls used: 0/60 · Total rewards: 0. Best-arm probabilities are computed by numerical integration of the posteriors.

Challenge: Thompson's trialGather at least 95% posterior probability that one of three slot machines is best, within 60 pulls.+50 XP

Match · Expression ↔ Meaning

Bayesian estimates

+20 XP
arg⁡max⁡θp(D∣θ)\arg\max_\theta p(D\mid\theta)
arg⁡max⁡θp(θ∣D)\arg\max_\theta p(\theta\mid D)
E[θ∣D]\mathbb E[\theta\mid D]

Options

Match · Expression ↔ Meaning

Beta updates

+20 XP
α←α+1\alpha\leftarrow\alpha+1
β←β+1\beta\leftarrow\beta+1
α/(α+β)\alpha/(\alpha+\beta)

Options

Proof puzzle

Beta-Bernoulli conjugacy

+25 XP

Claim

Update a Beta(alpha,beta) prior after s successes and f failures.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Gaussian prior becomes L2

+35 XP

Claim

Show MAP with w∼N(0,τ2I)w\sim\mathcal N(0,\tau^2I) minimises NLL plus ∥w∥2/(2τ2)\|w\|^2/(2\tau^2).

Preview

Your typeset proof appears here.

Coding problems

Problem 22·Warm-up

Update a prior

+20 XP

A Beta(2,2) prior sees 18 successes and 12 failures. Compute the posterior predictive probability of one more success as a reduced fraction.

An exact integer (or a fraction like 7/12)

Problem 23·Standard

Combine Gaussian evidence

+35 XP

The prior on μ is normal with mean 0 and variance 4. Four observations have known noise variance 1 and sample mean 3. Compute the posterior mean as a reduced fraction.

An exact integer (or a fraction like 7/12)

Problem 24·Challenge

Predict a run of successes

+50 XP

After observing data, θ has posterior Beta(3,2). Compute the probability that the next four conditionally independent Bernoulli trials all succeed, integrating over θ. Submit a reduced fraction.

An exact integer (or a fraction like 7/12)

Key takeaways

  • A posterior combines a stated prior and likelihood.
  • MLE, MAP and posterior mean answer different questions.
  • Posterior prediction integrates parameter uncertainty.
  • Thompson sampling uses posterior uncertainty to guide exploration.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

A Beta(2,3) prior sees three successes and one failure. What is the new alpha parameter?

Question 2 of 6 +20 XP

What is the mean of Beta(5,4)?

Question 3 of 6 +20 XP

A Gaussian prior on weights corresponds to what MAP penalty?

Question 4 of 6 +20 XP

What does the posterior predictive do?

Question 5 of 6 +20 XP

Beta(3,5) has an interior mode. What is it?

Question 6 of 6 +20 XP

Can sixty pulls guarantee identifying the best of any three Bernoulli arms with 95% posterior probability?

End of the chamber

Clear this chamber

+60 XPConjugate PriorMAP EstimatePosterior PredictiveThompson Sampling