Is Model A Really Better Than B?
Determining whether Model A's 0.5% offline metric gain over Model B is statistically significant or random sampling noise.
Why Point Estimates Are Misleading
Model A scores AUC = 0.885; Model B scores AUC = 0.880.
Is Model A genuinely superior, or did it get lucky on a favorable test set split?
95% Confidence Intervals of Metric Difference ΔAUC
Scenario 1: [==== (0.001 to 0.009) ====] ──► Statistically Significant (Interval excludes 0!)
Scenario 2: [====== (-0.003 to 0.013) ======] ──► NOT Statistically Significant (Interval crosses 0!)
If the 95% confidence interval of crosses zero, Model A is NOT provably better than Model B.
Three Rigorous Model Comparison Tests
1. McNemar's Test (Paired Binary Predictions)
Fast, non-parametric test operating on a contingency table of paired predictions on the same fixed test set:
Model B Correct Model B Incorrect
Model A Correct n_00 n_01 (A right, B wrong)
Model A Incorrect n_10 (B right, A wrong) n_11
McNemar Test Statistic (Chi-Square with 1 degree of freedom):
If (), the disagreement rates are statistically significant!
2. 5x2-Fold Cross-Validation Paired t-Test (Dietterich, 1998)
Corrects for overlapping fold correlations:
- Run 2-fold cross-validation 5 times (total 10 folds).
- Compute variance strictly within each 2-fold replication to avoid cross-replication correlation bugs.
3. Bootstrap Resampling Confidence Intervals (Gold Standard)
No model re-training required!
- Given saved test predictions and for test samples.
- For :
- Sample indices with replacement from test set.
- Compute .
- Sort 1,000 values. 95% CI is .
Say this out loud
"To prove Model A is statistically better than Model B, we construct confidence intervals around metric differences Δ = Metric_A - Metric_B. Standard paired t-tests over 10-fold CV fail due to training set overlap correlations. We use McNemar's test for paired classification outputs, or Bootstrap Resampling over saved test predictions to construct 95% confidence intervals around Δ metric gains."
Follow-ups to expect
- What is the Bonferroni Correction for multi-model comparison? If comparing 10 candidate models against a baseline, running 10 hypothesis tests at yields a high multi-testing false positive rate (). Bonferroni divides significance threshold by number of tests: .
- How large does a test set need to be to detect a 0.5% AUC lift? Depends on variance. Using power calculations for paired proportions, detecting a 0.5% difference with 80% power typically requires paired test samples.
Check yourself
Why does running a standard Paired Student's t-test across 10 folds of standard 10-fold cross-validation yield an artificially inflated Type I error rate?