Skip to content
AriadneTechnology

The Inner Ring · Chamber 9 of 9

Intervals, Tests and Honest Evaluation

Confidence intervals, p-values, permutation tests, the bootstrap and multiple comparisons: how to tell if one model is really better.

45 min 60 XP + 9 questions + 1 challengeMathVideoPapersProofsCodeLab

In this chamber you will

  • Build confidence intervals and say exactly what 95% means
  • Compute p-values with exact, permutation and normal-approximation tests
  • Compare two classifiers on the same test set with McNemar's test
  • Quantify uncertainty with the bootstrap, and control for multiple comparisons
DiscoverLearnRead beyondPapers & lecturesYour turn

Two classifiers disagree on a shared test set. Their accuracy difference is paired evidence: the same examples challenge both models. A valid comparison must preserve that pairing and account for how the models and evaluation were selected.

Spotted in the wild

χ2=(∣b−c∣−1)2b+c\chi^2=\frac{(|b-c|-1)^2}{b+c}
Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms
DiscoverLearnRead beyondPapers & lecturesYour turn
Symbols for this chamber
  • α\alpha“significance level”
    Target Type I error probability for a valid test.
  • H0H_0“null hypothesis”
    Model under which the test’s reference distribution is computed.
  • p-valuep\text{-value}“p-value”
    Null probability of a result at least as extreme as observed.
  • Xˉ±z SE\bar X\pm z\,\mathrm{SE}“mean plus or minus critical value times standard error”
    A normal-based interval formula under its assumptions.
  • b,cb,c“discordant counts”
    Paired examples where exactly one classifier is correct.
  • α/m\alpha/m“Bonferroni threshold”
    Per-test level controlling familywise error across m tests.

An interval is a procedure

For iid normal data with known standard deviation σ, Xˉ±1.96σ/n\bar X\pm1.96\sigma/\sqrt n has about 95% coverage. If σ is estimated from a normal sample, use the sample standard deviation and a Student t critical value with n−1 degrees of freedom. A normal approximation with an estimated standard error is an asymptotic method, not an exact small-sample result.

The frequentist statement concerns repeated random intervals covering a fixed parameter. After seeing one interval, the parameter is still fixed; assigning it a 95% posterior probability requires a Bayesian model. Reporting the interval width also shows the scale of uncertainty, which a binary significance label hides.

Quick check +20 XP

What does 95% frequentist coverage mean?

A p-value is conditional on a null model

Choose a null hypothesis, a test statistic and which outcomes count as at least as extreme. The p-value is the probability under the null of such an outcome. It is not the probability the null is true, nor the probability that chance caused the result.

A significance level α controls the test’s Type I error under its assumptions. Power is the probability of rejection under a specified alternative. A nonsignificant result may reflect little evidence or low power; it does not establish equivalence. A tiny p-value with a huge sample can accompany an unimportant effect.

Quick check +20 XP

A p-value is which probability?

Preserve the pair

For two fixed classifiers on the same independently sampled test examples, count b examples only A gets right and c only B gets right. Under equal discordant win probabilities and conditional on n=b+c, the number of A-only wins is Binomial(n,1/2).

A common two-sided exact McNemar p-value is min⁡(1,2P(B≤min⁡(b,c)))\min(1,2P(B\le\min(b,c))). If there are no disagreements, there is no evidence of a difference and we use p=1. The large-count approximation in the opening equation can be useful, but the exact calculation avoids its approximation error.

For paired numeric measurements, a sign-flip permutation test needs exchangeability under the null. Shuffling model outcomes independently across subjects breaks the pairing and changes the question.

Bootstrap the unit you actually sampled

A nonparametric bootstrap draws n observations with replacement from the observed n and recomputes the statistic. Repeating this produces an empirical distribution for the statistic; percentile endpoints are one possible interval method. They are approximate and can perform poorly for small samples, severe bias or nonregular statistics.

For paired model differences, resample whole paired examples. For clustered data, resample clusters when that matches the sampling design. The bootstrap cannot recover a missing population or undo a biased dataset.

Multiple comparisons and selection

If m valid null tests each use level α/m, the union bound limits the chance of any false rejection to at most α. This Bonferroni correction does not require independence. With independent tests each at α, the chance of at least one false rejection is 1−(1−α)m1-(1-\alpha)^m.

Repeatedly checking results and stopping at the first p<0.05p<0.05 also changes the procedure’s false-positive rate. Predefine a stopping rule or use a method designed for sequential testing. Tune models on validation data, then evaluate the chosen procedure on held-out data. Report effect sizes, uncertainty, seeds and the comparisons considered.

Quick check +20 XP

You test 20 hypotheses and want Bonferroni familywise level 0.05. What is the per-test threshold?

DiscoverLearnRead beyondPapers & lecturesYour turn

Read beyond

Book · free online · ~20 min

Introduction to Probability, Statistics, and Random Processes

Hossein Pishro-Nik · Chapter 8: estimation and testing

Compare frequentist coverage with Bayesian posterior probability.

Book · free online · ~20 min

OpenIntro Statistics

Diez, Çetinkaya-Rundel & Barr · Inference for means, proportions and paired data

Read the assumptions next to each interval and test.

Book · free online · ~20 min

Mathematics for Machine Learning

Deisenroth, Faisal & Ong · Chapter 8: model selection

Separate training, validation decisions and final evaluation.

DiscoverLearnRead beyondPapers & lecturesYour turn

Read the equation in context

Approximate Statistical Tests for Comparing Supervised Classification Learning AlgorithmsThomas G. Dietterich · 1998

Dietterich compares tests for classifier evaluation and discusses McNemar’s test for paired predictions. The approximate statistic uses discordant counts; for a small count, an exact binomial calculation is preferable. Comparing two already-trained models on fixed test data also addresses a narrower question than comparing entire training algorithms.

Decode the paper · Continuity-corrected McNemar statistic for discordant paired outcomes, for b+c>0

Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms

Thomas G. Dietterich · 1998

+25 XP
χ2=(∣b−c∣−1)2b+c\chi^2=\frac{(|b-c|-1)^2}{b+c}

Dietterich compares tests for classifier evaluation and discusses McNemar’s test for paired predictions. The approximate statistic uses discordant counts; for a small count, an exact binomial calculation is preferable. Comparing two already-trained models on fixed test data also addresses a narrower question than comparing entire training algorithms.

bb
cc
b+cb+c
χ2\chi^2

Options

Confidence Intervals, Clearly ExplainedStatQuest with Josh Starmer
DiscoverLearnRead beyondPapers & lecturesYour turn

Your turn

Repeat an interval-building procedure on simulated data with a known mean. Find an interval that is too narrow, repair its standard error, and compare the long-run coverage.

Interactive lab

Interval auditor

Simulate independent normal observations with true mean 1 and known standard deviation 2. Build one thousand nominal 95% intervals. Diagnose the narrow formula, then repair it.

Challenge: Interval auditorCatch a 95% interval that covers far less than 95% of the time, then repair it.+40 XP

Match · Expression ↔ Meaning

Error types

+20 XP
Reject a true null
Fail to reject a specified false null
Reject under a specified alternative

Options

Match · Expression ↔ Meaning

Evaluation procedures

+20 XP
Resample whole paired observations
Use only discordant correctness counts
Test each hypothesis at α/m

Options

Proof puzzle

Bonferroni from the union bound

+25 XP

Claim

If each of m tests has Type I error at most α/m, prove familywise error is at most α.

Tap lines in the order they should appear. Tap a line in your proof to send it back.

Your proof

  1. Pick the first line below.

Available lines

Prove it yourself

Coverage of a Gaussian interval

+35 XP

Claim

For iid normal data with known σ, derive coverage P(μ∈[Xˉ−zσ/n,Xˉ+zσ/n])=P(∣Z∣≤z)P(\mu\in[\bar X-z\sigma/\sqrt n,\bar X+z\sigma/\sqrt n])=P(|Z|\le z) for standard normal Z.

Preview

Your typeset proof appears here.

Coding problems

Problem 25·Warm-up

Exact paired evidence

+20 XP

Two fixed classifiers disagree on 10 test examples: A alone is right on 8, B alone on 2. Use the two-sided exact rule min⁡(1,2P(Binomial(10,1/2)≤2))\min(1,2P(Binomial(10,1/2)\le2)). Submit a reduced fraction.

An exact integer (or a fraction like 7/12)

Problem 26·Standard

Multiple independent chances

+35 XP

For twenty independent true-null tests, each with false-positive probability 0.05, compute the probability of at least one false positive to 6 decimal places.

A number, rounded to 6 decimal places

Problem 27·Challenge

An exact bootstrap distribution

+50 XP

The observed sample is (0,1,2). Enumerate all 27 ordered bootstrap resamples of length 3, each drawn with replacement from these three observations. What is the variance of their sample means? Submit a reduced fraction.

An exact integer (or a fraction like 7/12)

Key takeaways

  • Coverage describes a repeated-sampling procedure; credible intervals describe a posterior.
  • P-values condition on a null model and an analysis protocol.
  • Paired evaluation and resampling must preserve the data’s dependence structure.
  • Multiple comparisons, model selection and stopping rules affect the evidence.

Checkpoint

Prove it to the labyrinth

Answer every question to clear this chamber. First-try answers earn the most XP.

0/6
Question 1 of 6 +20 XP

Paired classifiers have b=7 and c=3. How many observations inform the conditional McNemar calculation?

Question 2 of 6 +20 XP

How should a paired bootstrap resample two classifiers’ test results?

Question 3 of 6 +20 XP

Does a nonsignificant test establish equality?

Question 4 of 6 +20 XP

Known sigma=2, n=100. What is the half-width of a normal interval using 1.96?

Question 5 of 6 +20 XP

Does Bonferroni control require independent tests?

Question 6 of 6 +20 XP

Which selection practice invalidates an ordinary fixed-sample p-value interpretation?

End of the chamber

Clear this chamber

+60 XPConfidence IntervalP-ValueBootstrapMultiple Comparisons