Chapter 16 of 60

Debugging Evaluation

Concepts

CHAPTER 16 β€” DEBUGGING EVALUATION

PART III β€” Debugging Interactive and Numerical AI

PURPOSE

Closes Part III on the honest-training dishonest-score case (0.97 validation vs 0.55 hand-written intent slice, same weights): the evaluation as a calibratable instrument perturbed without touching the model β€” seed/split swaps, fresh slices, metric swaps β€” each multi-seeded with pre-written score movements.

CENTRAL QUESTION

When the score passes and the intent fails, what ordered probe convicts the evaluation β€” before anyone re-trains anything?

UNIQUE CLAIM

An evaluation is a measurement device with its own error model debugged by perturbing the device (never the weights): re-audit split integrity (H1 leakage/split-rot β€” the Ch13 concept re-pointed at the instrument, not repeated), slice the score against baselines plus a CheckList-built intent slice (MFT/INV/DIR β€” H2 metric-vs-intent lead), then run seed-swap / decontaminated-fresh-slice / metric-swap perturbations β€” where H1 needs a large collapse toward baseline/intent because a 3–15% fresh-set drop is the intrinsic cost of an independent draw (Recht), not leakage; conviction requires the predicted movement pattern, all else is UNKNOWN.

DEBUGGING OBJECT

Evidence as instrument readings β€” headline metric/n/seed + weights/data hashes, overlap % + split provenance, per-class/cohort/date slices + majority baseline, three perturbation series; the 11%-overlap + fresh-0.58 + balanced-0.71 + seeds-0.93–0.97 H1+H2 joint conviction (H3 luck exonerated).

CONCEPTS INTRODUCED

Eval-perturbation probe (auditβ†’slicesβ†’three perturbations, weights frozen); fresh-slice calibration (large-collapse threshold); intent slice as ad-hoc CheckList (MFT minimum-functionality / INV invariance / DIR directional); metric validity as an argued task-scoped claim (not inherited); repaired-eval-as-contract (regenerated split + aligned metric + slice floors + tolerance + hashes); fourth disease (wrong test labels); “maximum wearing a mean’s clothes” (best-of-N as headline fraud).

CONCEPTS DEVELOPED / REUSED

Leakage concept from Ch13 (same diagnostic, different object β€” training-input vs eval-instrument, audit re-run because the object changed); behavioral testing stance (CheckList ~3x bugs β€” the instrument’s systematic form); single-run fallacy from Ch1/Ch12 (checkpoint cherry-picking, global-only reporting, n-as-validity); contract pinning from Ch8 (repaired eval as CI contract).

PREREQUISITES

Ch13 (leakage/splits), Ch15 (honest training), Ch12 (seeds/spreads), Ch8 (contracts).

LOCAL INVARIANTS

Freeze headline (metric/n/seed/weights/data) before perturbing; re-audit the split (overlap/ordering/provenance vs preprocessing history) first; slice before theorizing (cohorts + baseline + intent sample); pre-write numeric movement FORECASTs per hypothesis; all stochastic perturbations β‰₯3 seeds; repair-and-pin before any retraining.

FAILURE MODES

Retrain-first reflex (optimizing memorization, calling it progress); n-as-validity (50k samples measuring the wrong thing precisely); global-only reporting (0.97 averaging over a 0.55 that matters); checkpoint cherry-picking; metric inertia (accuracy/BLEU/exact-match by habit); fresh-slice theater (held-out drawn by the same leaking pipeline).

DIAGNOSTIC METHOD

  1. Freeze headline. 2. Re-audit split. 3. Slice + baseline + intent sample. 4. Seed-swap Γ—3 / fresh-slice / metric-swap with FORECASTs. 5. Pattern-match to H1/H2/H3 (+ label-error H4); repair (regen split + purge + aligned metric + floors + tolerance) into CI.

RESEARCH-DERIVED IDEAS

Ribeiro et al. ACL 2020 Best Paper CheckList (held-out accuracy overestimates; capability taxonomy Γ— MFT/INV/DIR; users found ~3x bugs incl. deployed commercial β€” NLP-bounded); Recht et al. ICML 2019 (replicated ImageNet/CIFAR collection: 11–14% / 3–15% drops, mostly harder-images not adaptivity, rankings held β€” vision-bounded; sets the H1 bar); Reiter 2018 (BLEU OK for system-level MT, not sentences/non-MT where still used); Dwork et al. Science 2015 reusable holdout (adaptive queries against a holdout degrade its validity in theory; DP-based safe-reuse mechanism) vs Roelofs et al. NeurIPS 2019 meta-analysis (112 Kaggle comps: little adaptive overfitting in practice, public/private scores track closely β€” EXCEPT with non-i.i.d. or small splits) β†’ adaptive overfitting concentrates exactly where H1 points; H3 checkpoint cherry-picking = adaptive querying in miniature; Northcutt et al. NeurIPS D&B 2021 (~3.4% label errors across 10 benchmarks; corrections move rankings); LLM benchmark contamination (pretraining scrape) = split-rot upstream of you β€” Ch50 owns it.

EXPERIMENT / LAB

Lab 16 (PROPOSED): suspect high score (or injected 10% trainβ†’val leak / 95-5 accuracy game / best-of-5 cherry-pick), weights/code/data frozen, H1/H2/H3 numeric FORECASTs (“fresh drops β‰₯0.25; seed moves <0.03”), full five-row perturbation table. H-structure: independent var = the evaluation (split/seed/metric/slice); controls = weights/code/data. Success = pattern-matching table + repaired-eval spec; retrain-without-table is not completion.

COMPANION TOOL

Evaluation Comparator β€” accepts: split audit + slice table + baseline + three perturbation series with FORECASTs + metric/tolerance spec. Can-establish: whether the headline measures leakage, wrong objective, or luck β€” which disease, under examined splits/metrics only. Cannot-establish: why weights fail the slice (interior β€” Part IV), training-data honesty beyond the audit (Ch13), future validity (drift rots it); never seed-agreement/high-score/symptom-improvement verdicts.

PREVENTION ARTIFACT

Repaired-eval specification (regen script + purged IDs + pinned metric + slice floors + tolerance + hashes) enforced in CI with scheduled re-verification.

READER OUTCOME

Reader can kill a retraining cycle with a perturbation table that names the eval disease β€” testable via Lab 16’s five-row record; also carries Part III’s closing map (Ch10 order β†’ 11 state β†’ 12 pins β†’ 13 data β†’ 14 tensors β†’ 15 training β†’ 16 exam).

DEPENDENCIES

Ch13, Ch15, Ch12, Ch8.

FORWARD BRIDGE

Part IV “Debugging Models” (Ch17 opacity stance) β€” inherits the certified-failure handoff: the model fails the intent slice repeatably and no leak/metric/seed explains it, so the suspect moves behind glass (inputs/outputs observable, mechanism not) with distributions replacing points.

EVIDENCE / RESEARCH REQUIREMENTS

0.97-vs-0.55 illustration constructed; Recht 3–15% intrinsic-drop calibration mandatory before H1 claims; CheckList NLP / Recht vision / Reiter BLEU scopes kept; fresh slices need independent sourcing or purge provenance.

ANTI-CLAIMS / LIMITS

One record convicts one eval disease under one weights/data/metric revision; explains nothing about weights interior; certifies no model; freezes no future validity; UNKNOWN where single-seeded or fresh slice shares the pipeline.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part III β€” Debugging Interactive and Numerical AI

0.97 and broken

Training is honest now (Chapter 15): gradients flow, the overfit gate passes, the curve descends on clean data through asserted handoffs. Validation accuracy: 0.97. The demo delights. Then the engineer tries ten hand-written cases reflecting actual intent β€” paraphrased requests, edge phrasings, the minority class the customer cares about:

validation accuracy: 0.97 (MEASUREMENT, n=2000, seed 42)
intent slice (hand-written, n=40): 0.55 (MEASUREMENT, same weights, same code)

The eval says pass. The intent says fail. Both numbers are real measurements of the same model.

OBSERVATION: 0.97 validation vs. 0.55 intent-slice; gap reproduces across reruns with fixed seed. HYPOTHESIS H1 (leakage/split rot): the validation set shares answers with training (overlap, time-travel, stale split regenerated after preprocessing changed) β€” the 0.97 measures memorization. HYPOTHESIS H2 (metric-vs-intent mismatch): the metric rewards behavior the intent does not want (accuracy on a 95%-majority split, BLEU on paraphrases, exact-match on equivalent answers) β€” the 0.97 measures the wrong thing correctly. HYPOTHESIS H3 (single-seed luck): the headline number is a favorable draw (seed, split sample, checkpoint cherry-pick) β€” re-draws regress toward the intent slice. INFERENCE: none yet β€” one validation number cannot separate memorization from mismeasurement from luck. Only perturbations of the evaluation (not the model) with pre-written score-movement predictions separate H1/H2/H3.

This chapter’s question: when the score passes and the intent fails, what ordered probe convicts the evaluation β€” before anyone re-trains anything?

Why “collect more validation data” fails first

The obvious move β€” enlarging the validation set by the same pipeline β€” fails because it replicates the pipeline’s defect at larger n. Four evaluation diseases hide behind high scores:

  1. Leakage and split rot. Train/validation ID overlap, near-duplicate paraphrases across the split boundary, or a split generated once and never regenerated while preprocessing, filtering, or labels changed underneath it. Chapter 13 diagnosed this same overlap as a data defect β€” a property of the training input. Here it is an instrument defect: the evaluation is the thing being calibrated, and a leaking split makes the instrument read high regardless of the model. Same concept, different debugging object; the audit is re-run because the object changed, not because it was skipped before. The split rots; the score stays green because it measures the rot’s consistency. For a hosted LLM the split can also rot upstream of you: benchmark items scraped into pretraining data are leakage you did not commit and cannot see in your own pipeline β€” Chapter 50 treats that contamination directly.
  2. Metric-vs-intent mismatch. Accuracy on imbalanced classes (0.97 by always predicting majority), ROUGE/BLEU rewarding surface overlap while intent needs factual equivalence, pass@k rewarding any-of-many while deployment takes the first. The metric is a contract (Chapter 8) β€” and this one was signed without reading.
  3. Slice blindness. A global 0.97 averaging over a 0.55 on the only slice that matters: the minority class, the new users, the adversarial phrasings, the post-cutoff dates. Aggregation is anesthesia.
  4. Seed and checkpoint luck. One seed, one split, best-of-20 checkpoints reported. The headline is the maximum of a distribution presented as its mean β€” single-run inference (book invariant) in evaluation clothes.

OPINION: an evaluation is a scientific instrument, and most teams calibrate it never. A debugger who trusts an uncalibrated instrument debugs the model for the instrument’s sins.

Two results anchor this. Ribeiro and colleagues showed that held-out accuracy routinely overestimates real capability, and that a structured set of hand-written behavioral tests surfaces bugs the aggregate score hides β€” in a user study, testers using their CheckList method found nearly three times as many bugs, including actionable failures in deployed commercial systems (Ribeiro et al., 2020). Recht and colleagues rebuilt the ImageNet and CIFAR-10 test sets by replicating the original collection process and found every model dropped 11–14% (ImageNet) or 3–15% (CIFAR-10) on the new sets β€” and traced most of that gap not to adaptive overfitting but to the new images being slightly harder (Recht et al., 2019). The lesson for the fresh-slice probe below: expect some drop on any genuinely independent slice; a leak shows up as a large collapse, not as a few points.

The mental model: evaluation is a measurement device with its own error model β€” and it debugs like any device: perturb the device, predict the reading’s movement, observe. Swap the seed, slice the data, stress the metric: a healthy eval moves as predicted; a diseased one reveals its disease by moving wrong (or not at all).

The method: the eval-perturbation probe

Ordered probe β€” the model weights never change; only the evaluation moves, one perturbation per run:

  1. Re-audit split integrity (H1). ID/n-gram overlap across train/validation, timestamp ordering, split-generation provenance (script hash + date vs. preprocessing-change dates). A stale or overlapping split convicts H1 before any perturbation runs. Record as MEASUREMENT.
  2. Slice the score (H2 lead). Report the metric per class / per cohort / per time bucket, plus a majority-class baseline and (where affordable) a 40-case hand-written intent slice. Prediction if H2: global stays high while the intent-critical slice collapses β€” the metric and the intent disagree by construction.
  3. Run the discriminating perturbations with pre-written score movements. (a) Seed/split swap: re-split with a new seed (or evaluate 3 seeds) β€” prediction if H3: score swings toward the intent slice (luck exonerated otherwise if stable); (b) Decontaminated slice: evaluate on a guaranteed-fresh held-out slice (post-cutoff, hand-written, or overlap-purged) β€” prediction if H1: the 0.97 collapses on the fresh slice while training accuracy holds; (c) Metric swap: score the same predictions with an intent-aligned metric (balanced accuracy / F1-minority / human pass on the slice) β€” prediction if H2: the headline drops while predictions are unchanged, proving the number was the metric’s, not the model’s. Because eval draws vary, each perturbation runs β‰₯3 seeds; single-perturbation verdicts are UNKNOWN.
  4. Convict the instrument, then repair it. Fresh split generation, overlap purge, intent-aligned metric pinned with a regression assertion β€” the evaluation becomes a Chapter 8 contract: inputs hashed, metric named, tolerance declared.
    flowchart TD
    A["re-audit the split: ID / n-gram overlap, timestamp ordering, split provenance vs preprocessing dates"] --> AD{"stale or overlapping split?"}
    AD -->|yes| H1a["H1 live: split rot inflates the headline"]
    AD -->|no| SL["slice the score: per class / cohort / time bucket + majority baseline + hand-written intent slice"]
    SL --> P["run 3 perturbations, weights frozen, >=3 seeds each"]
    P --> M{"which reading moved?"}
    M -->|"swings across seeds"| H3["H3: seed / checkpoint luck"]
    M -->|"large collapse only on the fresh slice"| H1["H1: leakage / split rot"]
    M -->|"drops only on metric swap / slices, predictions unchanged"| H2["H2: metric-vs-intent mismatch"]
    M -->|"only a few-point drop on the fresh slice"| OK["normal cost of an independent draw β€” not a conviction"]
  
# eval-perturbation probe (weights frozen; only the instrument moves)
seeds = [42, 43, 44]
for s in seeds:  # H3: seed swap β€” prediction if luck: spread covers the intent slice
    print("seed", s, evaluate(weights, resplit(seed=s), metric="accuracy"))
print("slices:", evaluate_by_slice(weights, val, by=["class", "cohort", "date_bucket"]))
print("fresh-slice:", evaluate(weights, fresh_heldout, metric="accuracy"))      # H1 probe
print("metric-swap:", evaluate(weights, val, metric="balanced_accuracy"))       # H2 probe
print("baseline:", majority_class_baseline(val))  # the score stupidity must beat
# Predictions pre-written: H1 collapses only on fresh-slice; H2 collapses only on
# metric-swap/slices; H3 swings across seeds. Any other pattern -> UNKNOWN, re-audit.

OBSERVATION (constructed illustration, not a measured run): overlap audit found 11% validation IDs in train; fresh post-cutoff slice scored 0.58 (vs. 0.97 validation); metric swap to balanced accuracy gave 0.71 with identical predictions; seed re-splits ranged 0.93–0.97 (H3 exonerated β€” luck is not the story). UPDATED BELIEF: H1 + H2 jointly supported β€” split rot inflates the headline, and accuracy hides minority-class failure. The model was never 0.97 at anything the customer wants. INFERENCE: the fix is split regeneration + overlap purge + balanced-accuracy contract with slice floors β€” zero retraining until the instrument is repaired. Training on a repaired eval may still fail; that is Chapter 15’s jurisdiction, now measurable.

Research lineage: the instrument has been studied

Behavioral testing formalizes the intent slice. CheckList structures hand-written tests along two axes β€” a taxonomy of capabilities (negation, coreference, robustness to typos, fairness) and three test types: Minimum Functionality Tests (does it get the obvious case right?), Invariance Tests (does the output stay the same under a meaning-preserving change?), and Directional Expectation Tests (does it move the right way under a meaning-changing one?) (Ribeiro et al., 2020). The 40-case intent slice in this chapter is an ad-hoc CheckList; building it with those three test types makes it systematic and turns each failure into a regression test.

Fresh test sets read lower even when nothing leaked. Recht and colleagues’ reconstruction study means the fresh-slice probe needs a calibrated expectation: a 3–15% drop is the normal cost of an independent draw, and models that were better on the original benchmark stayed better on the new one (the ranking held). An H1 conviction needs a collapse toward the majority baseline or the intent slice, not a routine few-point decline (Recht et al., 2019).

The instrument’s own answer key can be wrong. Northcutt, Athalye, and Mueller estimated an average of about 3.4% label errors across ten widely used benchmark test sets, and showed that on the corrected sets model rankings shift β€” a model that looked worse can be genuinely better once the test labels are right (Northcutt, Athalye & Mueller, 2021). This is a fifth eval disease alongside leakage, metric mismatch, slice blindness, and seed luck: the score disagrees with intent because the gold labels disagree with intent. Its tell is that the model’s “errors” on a hand-audited sample are disproportionately cases where the model is right and the label is wrong. The repair is a corrected slice with its own provenance, not a retrain.

Metric validity is a claim to be argued. Reiter’s structured review of BLEU concluded that it correlates acceptably with human judgment for system-level machine-translation comparison but not for individual outputs or for non-translation tasks where it is nonetheless routinely used (Reiter, 2018). The metric-swap probe is doing what the field often skips: checking whether the chosen number actually tracks the intent before trusting it.

Reusing the holdout is a real hazard that is usually mild β€” except where H1 lives. In theory, every model-selection decision made by looking at the validation score spends a little of that set’s validity; Dwork and colleagues formalized this and built a differential-privacy-based mechanism for reusing a holdout safely (Dwork et al., 2015). In practice the degradation is smaller than the theory’s worst case: Roelofs and colleagues examined 112 Kaggle competitions β€” thousands of teams repeatedly scoring against a public leaderboard, then ranked once on a private set β€” and found little adaptive overfitting, with the public and private scores tracking closely, except in competitions with pathologies like non-i.i.d. splits or small test sets (Roelofs et al., 2019). The synthesis for this chapter: adaptive overfitting concentrates exactly where H1 already points β€” a small or leaky split β€” so the split-integrity audit does more work than a generic “you’ve queried the holdout too many times” worry. H3’s checkpoint cherry-picking is the same hazard in miniature: best-of-20 is twenty adaptive queries against the eval.

Lab 16: eval-perturbation probe with predicted score movement

PROPOSED, not executed: no author-measured results are reported. The evidence this chapter requires is the reader’s own perturbation table.

Setup. Take any high score under suspicion (or inject disease: leak 10% of train into validation; report accuracy on a 95/5 split; cherry-pick the best of 5 seeds). Freeze weights, code, and data hashes β€” the evaluation is the independent variable; everything else is controlled.

Task.

  1. Write H1/H2/H3 with distinct numeric score-movement FORECASTs before perturbing (e.g., “H1: fresh slice drops β‰₯0.25; seed swap moves <0.03”).
  2. Re-audit split integrity and record the MEASUREMENT row (overlap %, split provenance). Then slice the score (per-class/cohort + baseline + hand-written intent sample where feasible).
  3. Execute all three perturbations (seed swap Γ—3, fresh-slice, metric swap), each β‰₯3 seeds where stochastic. Record OBSERVATION (scores verbatim) and UPDATED BELIEF per row. Conviction requires the predicted movement pattern; all other patterns are UNKNOWN with the next audit named.
  4. Repair the instrument on paper: regenerated-split script, purged IDs, pinned metric + slice floors + tolerance.
Probe Perturbation FORECAST OBSERVATION UPDATED BELIEF
audit overlap + provenance H1: overlap > ___% ___ H1 live/exonerated
slices per-class + baseline H2: slice β‰ͺ global ___ H2 live/exonerated
seeds 3 re-splits H3: swing β‰₯ ___ ___, ___, ___ H3 live/exonerated
fresh held-out slice H1: collapse β‰₯ ___ ___ H1 convicted/suspended
metric intent-aligned metric H2: drop with same preds ___ H2 convicted/suspended

Success criterion. A completed perturbation table matching one pre-written pattern plus the repaired-eval specification (split script + metric + floors + tolerance). A retrained model without this table is explicitly not completion β€” the instrument first, the engine second.

Companion tool: Evaluation Comparator

What it accepts: split-integrity audit, slice table + baseline, the three perturbation series with pre-written FORECASTs, and the pinned metric + tolerance specification. What it performs: it enforces probe order (audit β†’ slices β†’ perturbations), checks observed movements against FORECASTs, requires β‰₯3 seeds per stochastic perturbation, refuses an eval-health verdict while any row is UNKNOWN, and stamps the repaired evaluation with data/metric/seed hashes. What it can establish: whether the headline score measures leakage, the wrong objective, or luck β€” and which disease, under the examined splits and metrics only. What it cannot establish: model-internal causes (why weights fail the intent slice), data-training honesty beyond the audit (Chapter 13), or future validity β€” distributions drift and today’s repaired eval rots without scheduled re-verification. It never treats agreement across seeds, a high score, or downstream-symptom improvement as diagnosis. How its output changes your next action: an H1 conviction routes to split regeneration + purge + re-audit; H2 routes to metric replacement + slice floors as CI contracts; H3 routes to multi-seed reporting policy; a repaired-but-still-failing intent slice routes forward β€” out of Part III entirely β€” with the evaluation certified and the model interior now the suspect.

Paper form, sufficient for this chapter:

Headline (metric/n/seed): ___ / ___ / ___   Intent slice: ___ / n=___
Audit: overlap ___% | split script ___ (date ___) | preprocessing changes since? Y/N
Slices (class/cohort/date): ___   Baseline (majority): ___
Seed swap (3): ___ ___ ___ (spread ___)   Fresh slice: ___   Metric swap: ___
CONVICTION: H1 / H2 / H3 (circle; pattern-match line: ___)
REPAIRED EVAL: split script ___ | metric ___ | slice floors ___ | tol ___ | hashes ___

Where a software implementation does not yet exist in the reader’s stack, this record is the tool. The perturb-the-instrument discipline precedes any automation.

Reusable procedure: every suspicious score gets this before retraining

  1. Freeze the headline β€” metric, n, seed, weights hash, data hashes.
  2. Re-audit the split β€” overlap, ordering, provenance vs. preprocessing history.
  3. Slice before theorizing β€” per-cohort scores, baseline, hand-written intent sample.
  4. Perturb the instrument β€” seed swap, fresh slice, metric swap, FORECASTs pre-written.
  5. Repair and pin β€” regenerated split, aligned metric, slice floors, CI enforcement.

Failure modes

  • Retrain-first reflex. Tuning a model against a leaking eval optimizes memorization and calls it progress. The eval is the exam; a leaked answer key makes every student look brilliant.
  • n-as-validity. “n=50,000, so the score is solid.” Large n measures the wrong thing precisely β€” sample size never repairs leakage or metric mismatch.
  • Global-only reporting. One number hiding a collapsed minority slice. Aggregation without slicing is optics, not measurement.
  • Checkpoint cherry-picking. Best-of-20 reported without the distribution. The headline is a maximum wearing a mean’s clothes β€” report spread or UNKNOWN.
  • Metric inertia. Keeping accuracy/BLEU/exact-match because “the team always used it.” A metric is a contract with intent; unsigned contracts do not bind.
  • Fresh-slice theater. A “held-out” slice drawn by the same leaking pipeline. Fresh means independently sourced or overlap-purged with provenance β€” otherwise it is the same exam with a new cover.

Limits, per contract: one perturbation record convicts one eval disease under one weights/data/metric revision; it does not explain why the weights fail, does not certify the model, and does not freeze future validity. UNKNOWN where perturbations are single-seeded or the fresh slice shares the pipeline.

References

  • Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020, pp. 4902–4912. https://doi.org/10.18653/v1/2020.acl-main.442
  • Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet Classifiers Generalize to ImageNet? Proceedings of the 36th International Conference on Machine Learning (ICML), 2019, pp. 5389–5400. http://proceedings.mlr.press/v97/recht19a.html
  • Ehud Reiter. A Structured Review of the Validity of BLEU. Computational Linguistics 44(3), 2018, pp. 393–401. https://doi.org/10.1162/coli_a_00322
  • Curtis G. Northcutt, Anish Athalye, and Jonas Mueller. Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. NeurIPS Datasets and Benchmarks Track, 2021. https://arxiv.org/abs/2103.14749
  • Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth. The Reusable Holdout: Preserving Validity in Adaptive Data Analysis. Science 349(6248), 2015, pp. 636–638. https://doi.org/10.1126/science.aaa9375
  • Rebecca Roelofs, Vaishaal Shankar, Benjamin Recht, Sara Fridovich-Keil, Moritz Hardt, John Miller, and Ludwig Schmidt. A Meta-Analysis of Overfitting in Machine Learning. Advances in Neural Information Processing Systems 32 (NeurIPS), 2019. https://papers.nips.cc/paper/9117-a-meta-analysis-of-overfitting-in-machine-learning

Debugging Checklist

  • Headline frozen (metric, n, seed, weights + data hashes)?
  • Split re-audited (overlap %, ordering, provenance vs. preprocessing changes)?
  • Slice table + baseline + intent sample recorded before perturbing?
  • H1/H2/H3 with numeric movement FORECASTs written before perturbations?
  • All three perturbations run (seed Γ—3, fresh slice, metric swap)?
  • Repaired eval specified (split script + metric + floors + tolerance) for CI?
  • No retraining launched against an unrepaired instrument?

What This Chapter Established

  • Evaluation as a calibratable instrument: split-integrity audit β†’ slicing β†’ seed/fresh/metric perturbations with pre-written movements, each multi-seeded.
  • The eval-perturbation probe separating leakage/split-rot (H1), metric-vs-intent mismatch (H2), and seed luck (H3), demonstrated on the 0.97-vs-0.55 case convicted as H1+H2 β€” constructed illustration, no measured runs claimed.
  • Research grounding: held-out accuracy overestimates capability and behavioral testing (CheckList: MFT/INV/DIR) finds ~3x more bugs (Ribeiro et al.); a fresh test set drops 3–15% even without leakage, so H1 needs a large collapse (Recht et al.); benchmark test sets carry ~3.4% mean label errors whose correction reorders model rankings (Northcutt et al.) β€” a fifth eval disease, the gold labels disagreeing with intent; metric validity is task-specific and must be argued, not inherited (Reiter); adaptive holdout reuse degrades validity in theory (Dwork et al.) but is empirically mild except with small or non-i.i.d. splits (Roelofs et al. β€” 112 Kaggle competitions), which is exactly where H1 already points. The intent slice is an ad-hoc CheckList.
  • Lab 16 as a proposed perturbation record the reader executes; the Evaluation Comparator contract (accepts/performs/can-establish/cannot-establish/next-action).
  • What was NOT proved: anything about the model’s interior β€” weights, representations, mechanisms. Every verdict here concerns the instrument, never the engine.
  • Part III’s closing map: Chapter 10 ordered the program, 11 explained its state, 12 pinned its reproduction, 13 honored its data, 14 asserted its tensors, 15 triaged its training, this chapter calibrated its exam. Interactive and numerical AI, end to end.

Next

The instrument is repaired and it reports, honestly: the model fails the intent slice. No leak explains it. No metric swap rescues it. No seed relitigates it. The evaluation has done everything an evaluation can do β€” it says that the model fails, on which slice, by how much, repeatably.

Notice what “repeatably” has come to mean. Chapters 12, 15, and 16 each stopped trusting a single run: reproduction became three cold runs, a training verdict needed a second seed, an eval verdict needed three. Deterministic debugging asked what value or state produced this failure and expected one answer. From here the question changes shape: what distribution of behavior appears under controlled, repeated conditions β€” because the answer is now a spread, not a point.

What the evaluation still cannot say is what inside the model fails: which input distinctions it ignores, which context steers it, which sampling draw doomed the answer. The failure is behind glass β€” observable only through inputs and outputs, never by inspection. Part IV opens that glass, and the two facts compound: less of the mechanism is visible, so controlled behavioral distributions become the evidence that replaces inspection.