An email containing “prize” is evidence for spam, but the strength of that evidence depends on how often the word occurs in both spam and legitimate mail. Bayes’ rule combines that comparison with the base rate.
Spotted in the wild
- “probability of H given E”Posterior belief after observing E.
- “probability of E given H”Likelihood of the evidence under H.
- “evidence probability”Normalising probability across all hypotheses.
- “odds of H”Probability divided by its complement.
- “log odds”An additive scale for evidence.
- “X independent of Y given C”The conditional factorisation used by naive Bayes.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “probability of H given E” | Posterior belief after observing E. | ||
| “probability of E given H” | Likelihood of the evidence under H. | ||
| “evidence probability” | Normalising probability across all hypotheses. | ||
| “odds of H” | Probability divided by its complement. | ||
| “log odds” | An additive scale for evidence. | ||
| “X independent of Y given C” | The conditional factorisation used by naive Bayes. |
Partition the possibilities
For disjoint hypotheses covering all possibilities, the law of total probability is
Bayes’ rule reverses a conditional: when . The denominator makes the posterior probabilities sum to one.
Suppose 20% of mail is spam. A word occurs in 40% of spam and 5% of legitimate mail. Its overall frequency is . Among messages containing it, the spam probability is .
Spam has prior probability 0.2. A word occurs with probability 0.4 in spam and 0.05 in ham. What is the posterior spam probability given the word?
Evidence multiplies odds
For a binary hypothesis H, odds are . Divide Bayes’ rule for H by the rule for its complement:
The last factor is the likelihood ratio. A ratio above one favours H; below one favours its complement. Odds return to probability by . Log-odds turn products into sums, reducing numerical underflow from many tiny probabilities.
Prior spam odds of , followed by likelihood ratios 8 and 3, become odds 6 and probability . Forgetting the prior gives a different answer.
Prior odds are 1/4 and evidence has likelihood ratio 8. What are the posterior odds?
Conditional independence is an assumption
Naive Bayes assumes that features are independent given the class. It does not assume they are independent without conditioning. Words such as “prize” and “winner” may remain correlated even within spam, so multiplying their evidence can be overconfident.
For binary presence features, a full Bernoulli model includes both present and absent words: . A multinomial word-count model instead uses token frequencies. Our lab uses a deliberately small presence-only evidence model, stated explicitly so its probabilities have a clear interpretation.
Zero estimated probabilities would erase a whole class likelihood. Additive smoothing uses positive pseudocounts; it changes the estimate rather than proving an unseen event impossible. Binary naive Bayes yields a log-odds expression linear in its features, resembling logistic regression. Naive Bayes fits class-conditional distributions; logistic regression fits the conditional class probability directly.
The observation protocol changes the likelihood
In standard Monty Hall, you select one of three doors. A host who knows the prize location always opens another door showing a goat and offers a switch. Switching wins exactly when your initial choice was wrong, with probability .
A host who opens another door uniformly at random without knowing the prize creates a different experiment. Conditional on revealing a goat, your original door and the other closed door each have probability . The observation “a goat appeared” has different likelihoods under the two protocols. Write those likelihoods before importing a familiar answer.
Why does the host protocol matter in Monty Hall?
Read beyond
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapter 1: conditional probability and independence
Calculate a posterior by first finding the evidence probability.
Book · free online · ~20 min
Introduction to ProbabilityBlitzstein & Hwang · Conditional probability lectures and book chapter
Compare the host protocols in versions of Monty Hall.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Section 6.3: sum rule, product rule and Bayes’ theorem
Translate between joint, conditional and marginal probabilities.
Read the equation in context
A Bayesian Approach to Filtering Junk E-MailMehran Sahami, Susan Dumais, David Heckerman & Eric Horvitz · 1998The paper treats junk-mail filtering as a probabilistic decision problem with different costs for different mistakes. The displayed factorisation is the naive Bayes modelling assumption. A posterior probability and the action chosen from it are separate: a costly false positive can justify a higher spam threshold.
Decode the paper · Naive Bayes classifier, restated with C for class and x_j for observed features
A Bayesian Approach to Filtering Junk E-MailMehran Sahami, Susan Dumais, David Heckerman & Eric Horvitz · 1998
The paper treats junk-mail filtering as a probabilistic decision problem with different costs for different mistakes. The displayed factorisation is the naive Bayes modelling assumption. A posterior probability and the action chosen from it are separate: a costly false positive can justify a higher spam threshold.
Options
Your turn
Build two three-word messages from the vocabulary. Use likelihood ratios to make one strongly favour spam and the other strongly favour legitimate mail.
Interactive lab
Tip the filter
| Word | P(word | spam) | P(word | ham) | Likelihood ratio |
|---|---|---|---|
| 0.4 | 0.02 | 20.00 | |
| 0.3 | 0.01 | 30.00 | |
| 0.5 | 0.05 | 10.00 | |
| 0.02 | 0.4 | 0.05 | |
| 0.01 | 0.3 | 0.03 | |
| 0.05 | 0.5 | 0.10 |
Message: Choose words above
Posterior spam probability
20.00%
Posterior log-odds
-1.386
Spam message: remaining · Ham message: remaining
Match · Expression ↔ Meaning
Bayesian quantities
Options
Match · Expression ↔ Meaning
Evidence scales
Options
Proof puzzle
Odds form of Bayes
Claim
Derive posterior odds = prior odds times likelihood ratio.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Normalise a posterior
Claim
For a finite partition H_i and positive P(E), prove .
Your typeset proof appears here.
Coding problems
Problem 1·Warm-up
A three-word posterior
Prior spam probability is . Three conditionally independent observed features have likelihood ratios 8, 1/2 and 3. Submit the posterior spam probability as a reduced fraction.
Problem 2·Standard
How much evidence is enough?
A rare hypothesis has prior probability 0.001. Each independent positive observation has likelihood ratio 3. Find the smallest count of observations that makes its posterior at least 0.95.
Problem 3·Challenge
Evidence strings
Consider all ordered length-10 strings of A and B. Under H, each symbol independently has P(A)=0.8; under its alternative, P(A)=0.2. Priors are equal. How many strings give posterior P(H | string)>0.99?
Key takeaways
- The base rate and the evidence likelihood both affect the posterior.
- Independent evidence multiplies odds and adds log-odds.
- Naive Bayes uses conditional independence as a modelling assumption.
- A sampling or observation protocol belongs in the likelihood.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Convert odds 3 to a probability.
Naive Bayes assumes which independence?
Two conditionally independent evidence items have likelihood ratios 2 and 5. What is their combined ratio?
What does a likelihood ratio below one mean?
What is additive smoothing used for?
In standard Monty Hall, what is the chance of winning by always switching?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Tip the filter (+30 XP)
- Bonus: Problem 1: A three-word posterior (+20 XP)
- Bonus: Problem 2: How much evidence is enough? (+35 XP)
- Bonus: Problem 3: Evidence strings (+50 XP)
- Bonus: Proof: Odds form of Bayes (+25 XP)
- Bonus: Proof: Normalise a posterior (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Bayesian quantities (+20 XP)
- Bonus: Match: Evidence scales (+20 XP)