Two classifiers disagree on a shared test set. Their accuracy difference is paired evidence: the same examples challenge both models. A valid comparison must preserve that pairing and account for how the models and evaluation were selected.
Spotted in the wild
- “significance level”Target Type I error probability for a valid test.
- “null hypothesis”Model under which the test’s reference distribution is computed.
- “p-value”Null probability of a result at least as extreme as observed.
- “mean plus or minus critical value times standard error”A normal-based interval formula under its assumptions.
- “discordant counts”Paired examples where exactly one classifier is correct.
- “Bonferroni threshold”Per-test level controlling familywise error across m tests.
| Symbol | Say it | Meaning | LaTeX |
|---|---|---|---|
| “significance level” | Target Type I error probability for a valid test. | ||
| “null hypothesis” | Model under which the test’s reference distribution is computed. | ||
| “p-value” | Null probability of a result at least as extreme as observed. | ||
| “mean plus or minus critical value times standard error” | A normal-based interval formula under its assumptions. | ||
| “discordant counts” | Paired examples where exactly one classifier is correct. | ||
| “Bonferroni threshold” | Per-test level controlling familywise error across m tests. |
An interval is a procedure
For iid normal data with known standard deviation σ, has about 95% coverage. If σ is estimated from a normal sample, use the sample standard deviation and a Student t critical value with n−1 degrees of freedom. A normal approximation with an estimated standard error is an asymptotic method, not an exact small-sample result.
The frequentist statement concerns repeated random intervals covering a fixed parameter. After seeing one interval, the parameter is still fixed; assigning it a 95% posterior probability requires a Bayesian model. Reporting the interval width also shows the scale of uncertainty, which a binary significance label hides.
What does 95% frequentist coverage mean?
A p-value is conditional on a null model
Choose a null hypothesis, a test statistic and which outcomes count as at least as extreme. The p-value is the probability under the null of such an outcome. It is not the probability the null is true, nor the probability that chance caused the result.
A significance level α controls the test’s Type I error under its assumptions. Power is the probability of rejection under a specified alternative. A nonsignificant result may reflect little evidence or low power; it does not establish equivalence. A tiny p-value with a huge sample can accompany an unimportant effect.
A p-value is which probability?
Preserve the pair
For two fixed classifiers on the same independently sampled test examples, count b examples only A gets right and c only B gets right. Under equal discordant win probabilities and conditional on n=b+c, the number of A-only wins is Binomial(n,1/2).
A common two-sided exact McNemar p-value is . If there are no disagreements, there is no evidence of a difference and we use p=1. The large-count approximation in the opening equation can be useful, but the exact calculation avoids its approximation error.
For paired numeric measurements, a sign-flip permutation test needs exchangeability under the null. Shuffling model outcomes independently across subjects breaks the pairing and changes the question.
Bootstrap the unit you actually sampled
A nonparametric bootstrap draws n observations with replacement from the observed n and recomputes the statistic. Repeating this produces an empirical distribution for the statistic; percentile endpoints are one possible interval method. They are approximate and can perform poorly for small samples, severe bias or nonregular statistics.
For paired model differences, resample whole paired examples. For clustered data, resample clusters when that matches the sampling design. The bootstrap cannot recover a missing population or undo a biased dataset.
Multiple comparisons and selection
If m valid null tests each use level α/m, the union bound limits the chance of any false rejection to at most α. This Bonferroni correction does not require independence. With independent tests each at α, the chance of at least one false rejection is .
Repeatedly checking results and stopping at the first also changes the procedure’s false-positive rate. Predefine a stopping rule or use a method designed for sequential testing. Tune models on validation data, then evaluate the chosen procedure on held-out data. Report effect sizes, uncertainty, seeds and the comparisons considered.
You test 20 hypotheses and want Bonferroni familywise level 0.05. What is the per-test threshold?
Read beyond
Book · free online · ~20 min
Introduction to Probability, Statistics, and Random ProcessesHossein Pishro-Nik · Chapter 8: estimation and testing
Compare frequentist coverage with Bayesian posterior probability.
Book · free online · ~20 min
OpenIntro StatisticsDiez, Çetinkaya-Rundel & Barr · Inference for means, proportions and paired data
Read the assumptions next to each interval and test.
Book · free online · ~20 min
Mathematics for Machine LearningDeisenroth, Faisal & Ong · Chapter 8: model selection
Separate training, validation decisions and final evaluation.
Read the equation in context
Approximate Statistical Tests for Comparing Supervised Classification Learning AlgorithmsThomas G. Dietterich · 1998Dietterich compares tests for classifier evaluation and discusses McNemar’s test for paired predictions. The approximate statistic uses discordant counts; for a small count, an exact binomial calculation is preferable. Comparing two already-trained models on fixed test data also addresses a narrower question than comparing entire training algorithms.
Decode the paper · Continuity-corrected McNemar statistic for discordant paired outcomes, for b+c>0
Approximate Statistical Tests for Comparing Supervised Classification Learning AlgorithmsThomas G. Dietterich · 1998
Dietterich compares tests for classifier evaluation and discusses McNemar’s test for paired predictions. The approximate statistic uses discordant counts; for a small count, an exact binomial calculation is preferable. Comparing two already-trained models on fixed test data also addresses a narrower question than comparing entire training algorithms.
Options
Your turn
Repeat an interval-building procedure on simulated data with a known mean. Find an interval that is too narrow, repair its standard error, and compare the long-run coverage.
Interactive lab
Interval auditor
Match · Expression ↔ Meaning
Error types
Options
Match · Expression ↔ Meaning
Evaluation procedures
Options
Proof puzzle
Bonferroni from the union bound
Claim
If each of m tests has Type I error at most α/m, prove familywise error is at most α.
Tap lines in the order they should appear. Tap a line in your proof to send it back.
Your proof
- Pick the first line below.
Available lines
Prove it yourself
Coverage of a Gaussian interval
Claim
For iid normal data with known σ, derive coverage for standard normal Z.
Your typeset proof appears here.
Coding problems
Problem 25·Warm-up
Exact paired evidence
Two fixed classifiers disagree on 10 test examples: A alone is right on 8, B alone on 2. Use the two-sided exact rule . Submit a reduced fraction.
Problem 26·Standard
Multiple independent chances
For twenty independent true-null tests, each with false-positive probability 0.05, compute the probability of at least one false positive to 6 decimal places.
Problem 27·Challenge
An exact bootstrap distribution
The observed sample is (0,1,2). Enumerate all 27 ordered bootstrap resamples of length 3, each drawn with replacement from these three observations. What is the variance of their sample means? Submit a reduced fraction.
Key takeaways
- Coverage describes a repeated-sampling procedure; credible intervals describe a posterior.
- P-values condition on a null model and an analysis protocol.
- Paired evaluation and resampling must preserve the data’s dependence structure.
- Multiple comparisons, model selection and stopping rules affect the evidence.
Checkpoint
Prove it to the labyrinth
Answer every question to clear this chamber. First-try answers earn the most XP.
Paired classifiers have b=7 and c=3. How many observations inform the conditional McNemar calculation?
How should a paired bootstrap resample two classifiers’ test results?
Does a nonsignificant test establish equality?
Known sigma=2, n=100. What is the half-width of a normal interval using 1.96?
Does Bonferroni control require independent tests?
Which selection practice invalidates an ordinary fixed-sample p-value interpretation?
End of the chamber
Clear this chamber
- Questions in this chamber (0/9 solved)Next unsolved
- Bonus: Interval auditor (+40 XP)
- Bonus: Problem 25: Exact paired evidence (+20 XP)
- Bonus: Problem 26: Multiple independent chances (+35 XP)
- Bonus: Problem 27: An exact bootstrap distribution (+50 XP)
- Bonus: Proof: Bonferroni from the union bound (+25 XP)
- Bonus: Proof: Coverage of a Gaussian interval (+35 XP)
- Bonus: Decode the paper (+25 XP)
- Bonus: Match: Error types (+20 XP)
- Bonus: Match: Evaluation procedures (+20 XP)