Skip to content
AriadneTechnology

The Outer Ring · Chamber 1 of 9

Bayes' Rule at Work: Odds and Evidence

Odds, likelihood ratios and evidence that adds up, from the Monty Hall problem to a naive Bayes spam filter.

40 min 50 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Use the law of total probability and Bayes' rule with many hypotheses
  • Update beliefs with odds, likelihood ratios and log-odds
  • Explain conditional independence and sequential updating
  • Build a naive Bayes classifier and compare it with logistic regression
DiscoverLearnRead beyondPapers & lecturesYour turn

An email containing “prize” is evidence for spam, but the strength of that evidence depends on how often the word occurs in both spam and legitimate mail. Bayes’ rule combines that comparison with the base rate.

Spotted in the wild

P(C∣x1,…,xd)∝P(C)∏j=1dP(xj∣C)P(C\mid x_1,\ldots,x_d)\propto P(C)\prod_{j=1}^dP(x_j\mid C)
A Bayesian Approach to Filtering Junk E-Mail
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • P(H∣E)P(H\mid E)“probability of H given E”
    Posterior belief after observing E.
  • P(E∣H)P(E\mid H)“probability of E given H”
    Likelihood of the evidence under H.
  • P(E)P(E)“evidence probability”
    Normalising probability across all hypotheses.
  • O(H)O(H)“odds of H”
    Probability divided by its complement.
  • log⁡O(H)\log O(H)“log odds”
    An additive scale for evidence.
  • X⊥Y∣CX\perp Y\mid C“X independent of Y given C”
    The conditional factorisation used by naive Bayes.

Partition the possibilities

For disjoint hypotheses HiH_i covering all possibilities, the law of total probability is

P(E)=∑iP(E∣Hi)P(Hi).P(E)=\sum_iP(E\mid H_i)P(H_i).

Bayes’ rule reverses a conditional: P(Hi∣E)=P(E∣Hi)P(Hi)/P(E)P(H_i\mid E)=P(E\mid H_i)P(H_i)/P(E) when P(E)>0P(E)>0. The denominator makes the posterior probabilities sum to one.

Suppose 20% of mail is spam. A word occurs in 40% of spam and 5% of legitimate mail. Its overall frequency is 0.4(0.2)+0.05(0.8)=0.120.4(0.2)+0.05(0.8)=0.12. Among messages containing it, the spam probability is 0.08/0.12=2/30.08/0.12=2/3.

Quick check +20 XP

Spam has prior probability 0.2. A word occurs with probability 0.4 in spam and 0.05 in ham. What is the posterior spam probability given the word?

Evidence multiplies odds

For a binary hypothesis H, odds are O(H)=P(H)/(1−P(H))O(H)=P(H)/(1-P(H)). Divide Bayes’ rule for H by the rule for its complement:

O(H∣E)=O(H)P(E∣H)P(E∣Hc).O(H\mid E)=O(H)\frac{P(E\mid H)}{P(E\mid H^c)}.

The last factor is the likelihood ratio. A ratio above one favours H; below one favours its complement. Odds return to probability by P=O/(1+O)P=O/(1+O). Log-odds turn products into sums, reducing numerical underflow from many tiny probabilities.

Prior spam odds of 1/41/4, followed by likelihood ratios 8 and 3, become odds 6 and probability 6/76/7. Forgetting the prior gives a different answer.

Quick check +20 XP

Prior odds are 1/4 and evidence has likelihood ratio 8. What are the posterior odds?

Conditional independence is an assumption

Naive Bayes assumes that features are independent given the class. It does not assume they are independent without conditioning. Words such as “prize” and “winner” may remain correlated even within spam, so multiplying their evidence can be overconfident.

For binary presence features, a full Bernoulli model includes both present and absent words: P(xj∣C)=pjCxj(1−pjC)1−xjP(x_j\mid C)=p_{jC}^{x_j}(1-p_{jC})^{1-x_j}. A multinomial word-count model instead uses token frequencies. Our lab uses a deliberately small presence-only evidence model, stated explicitly so its probabilities have a clear interpretation.

Zero estimated probabilities would erase a whole class likelihood. Additive smoothing uses positive pseudocounts; it changes the estimate rather than proving an unseen event impossible. Binary naive Bayes yields a log-odds expression linear in its features, resembling logistic regression. Naive Bayes fits class-conditional distributions; logistic regression fits the conditional class probability directly.

The observation protocol changes the likelihood

In standard Monty Hall, you select one of three doors. A host who knows the prize location always opens another door showing a goat and offers a switch. Switching wins exactly when your initial choice was wrong, with probability 2/32/3.

A host who opens another door uniformly at random without knowing the prize creates a different experiment. Conditional on revealing a goat, your original door and the other closed door each have probability 1/21/2. The observation “a goat appeared” has different likelihoods under the two protocols. Write those likelihoods before importing a familiar answer.

Quick check +20 XP

Why does the host protocol matter in Monty Hall?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapter 1: conditional probability and independence

Calculate a posterior by first finding the evidence probability.

Book · free online · ~20 min

Introduction to Probability

Blitzstein & Hwang · Conditional probability lectures and book chapter

Compare the host protocols in versions of Monty Hall.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Section 6.3: sum rule, product rule and Bayes’ theorem

Translate between joint, conditional and marginal probabilities.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

A Bayesian Approach to Filtering Junk E-MailMehran Sahami, Susan Dumais, David Heckerman & Eric Horvitz · 1998

The paper treats junk-mail filtering as a probabilistic decision problem with different costs for different mistakes. The displayed factorisation is the naive Bayes modelling assumption. A posterior probability and the action chosen from it are separate: a costly false positive can justify a higher spam threshold.

Decode the paper · Naive Bayes classifier, restated with C for class and x_j for observed features

A Bayesian Approach to Filtering Junk E-Mail

Mehran Sahami, Susan Dumais, David Heckerman & Eric Horvitz · 1998

+25 XP
P(C∣x1,…,xd)∝P(C)∏j=1dP(xj∣C)P(C\mid x_1,\ldots,x_d)\propto P(C)\prod_{j=1}^dP(x_j\mid C)

The paper treats junk-mail filtering as a probabilistic decision problem with different costs for different mistakes. The displayed factorisation is the naive Bayes modelling assumption. A posterior probability and the action chosen from it are separate: a costly false positive can justify a higher spam threshold.

CC
xjx_j
P(C)P(C)
∏jP(xj∣C)\prod_jP(x_j\mid C)

Options

Naive Bayes, Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Build two three-word messages from the vocabulary. Use likelihood ratios to make one strongly favour spam and the other strongly favour legitimate mail.

Interactive lab

Tip the filter

Prior spam probability is 20%. This toy model uses only selected words, at most once each, and assumes their evidence is independent given the class. It does not score absent words.
WordP(word | spam)P(word | ham)Likelihood ratio
0.40.0220.00
0.30.0130.00
0.50.0510.00
0.020.40.05
0.010.30.03
0.050.50.10

Message: Choose words above

Posterior spam probability

20.00%

Posterior log-odds

-1.386

Spam message: remaining · Ham message: remaining

Challenge: Tip the filterWrite a three-word message the filter calls spam with 99% confidence, and another it calls ham with 99%.+30 XP

Match · Expression ↔ Meaning

Bayesian quantities

+20 XP
P(H)P(H)
P(E∣H)P(E\mid H)
P(H∣E)P(H\mid E)

Options

Match · Expression ↔ Meaning

Evidence scales

+20 XP
Onew=OoldrO_{new}=O_{old}r
log⁡Onew=log⁡Oold+log⁡r\log O_{new}=\log O_{old}+\log r
p=O/(1+O)p=O/(1+O)

Options

Proof puzzle

Odds form of Bayes

+25 XP

Claim

Derive posterior odds = prior odds times likelihood ratio.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Normalise a posterior

+35 XP

Claim

For a finite partition H_i and positive P(E), prove ∑iP(Hi∣E)=1\sum_iP(H_i\mid E)=1.

Preview

Your typeset proof appears here.

Coding problems

Problem 1·Warm-up

A three-word posterior

+20 XP

Prior spam probability is 1/51/5. Three conditionally independent observed features have likelihood ratios 8, 1/2 and 3. Submit the posterior spam probability as a reduced fraction.

An exact integer (or a fraction like 7/12)

Problem 2·Standard

How much evidence is enough?

+35 XP

A rare hypothesis has prior probability 0.001. Each independent positive observation has likelihood ratio 3. Find the smallest count of observations that makes its posterior at least 0.95.

An exact integer (or a fraction like 7/12)

Problem 3·Challenge

Evidence strings

+50 XP

Consider all ordered length-10 strings of A and B. Under H, each symbol independently has P(A)=0.8; under its alternative, P(A)=0.2. Priors are equal. How many strings give posterior P(H | string)>0.99?

An exact integer (or a fraction like 7/12)

Key takeaways

  • The base rate and the evidence likelihood both affect the posterior.
  • Independent evidence multiplies odds and adds log-odds.
  • Naive Bayes uses conditional independence as a modelling assumption.
  • A sampling or observation protocol belongs in the likelihood.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

Convert odds 3 to a probability.

Question 2 of 6 +20 XP

Naive Bayes assumes which independence?

Question 3 of 6 +20 XP

Two conditionally independent evidence items have likelihood ratios 2 and 5. What is their combined ratio?

Question 4 of 6 +20 XP

What does a likelihood ratio below one mean?

Question 5 of 6 +20 XP

What is additive smoothing used for?

Question 6 of 6 +20 XP

In standard Monty Hall, what is the chance of winning by always switching?

End of the chamber

Clear this chamber

+50 XPTotal ProbabilityOddsLikelihood RatioNaive Bayes