Chapter 60 of 60

The Debugging AI Toolkit

Concepts

CHAPTER 60 — The Debugging AI Toolkit

PART X — The Debugging AI Playbook

PURPOSE

Closes the book by filing all sixty chapters into one workbench — drawer inventory with contracts, symptom→chapters→tools selection guide, single-spine composition — stating what the book proved within bounds and left provisional, with no new diagnoses.

CENTRAL QUESTION

What is the accumulated toolkit as one coherent kit — how is each instrument selected, what does each combination prove, and what does the book leave provisional?

UNIQUE CLAIM

Only this chapter provides budget-first composition: budget (10/60/full) selects procedure (Ch57/58/59), procedure selects drawers, every drawer runs the one spine (freeze → hypothesize → intervene single-variable → trial → artifact) — with compound symptoms split across drawers, gaps named as output, and generality bounded per drawer/symptom/budget. Book closing: no next chapter.

DEBUGGING OBJECT

“Wrong and slow” teammate compound symptom with no case sentence/freeze/owner: H1 selection failure (mapping missing) vs H2 instrument failure (uncomposable contracts) vs H3 proof failure (unstated establishes) — retired by the guide (triage → wrongness to Ch30–35 drawer with hash-diff queued, slowness to Ch55 ledger with probes queued), zero verdicts at triage.

CONCEPTS INTRODUCED (only genuinely new here)

  • Workbench drawer inventory — ten Parts as ten drawers, each entry a contract (accepts/performs/can/cannot/next-action), not a summary. Drawer map (reconciled with the enriched manuscript, audit §16.1 — already applied in 60-chapter.md): I Ch1–4 debugging from first principles; II Ch5–9 deterministic software; III Ch10–16 interactive and numerical AI; IV Ch17–23 models; V Ch24–29 AI-assisted development and research; VI Ch30–35 prompts/retrieval/hallucinations; VII Ch36–43 agents and trajectories; VIII Ch44–51 building the AI debugger; IX Ch52–56 production; X Ch57–60 playbook budgets. Each drawer entry names when to reach for it and what it outputs.
  • Selection guide (symptom → chapters → tools): wrong-single-case → Ch57 triage → Ch30-35; wrong-repeated → Ch58 → Ch2 bisection → Ch18-23 attribution; agent misbehaved → Ch36-43; hallucinated citation → Ch34; money/data → Ch56 → Ch52 → Ch53 → Ch54; slow/expensive → Ch55; “improved” unmeasured → Ch16 → Ch23; machine debugger proposed → Ch44 gates; full weight → Ch59; unsure → Ch57 route-never-verdict. RULE: budget first (57/58/59), then drawer.
  • READ-DO (per-chapter paper forms) vs DO-CONFIRM (Debugging Checklists) taxonomy; built-in obsolescence stated (deep-structure categories need re-indexing as model/retrieval/agent stacks change)
  • Composition contract: every drawer at every budget runs the one spine (freeze → hypothesize → intervene single-variable → trial → artifact); compound symptoms split across drawers; gaps named as output; generality bounded per drawer/symptom/budget; no new diagnoses in the closing.

CONCEPTS DEVELOPED / REUSED (with source chapter)

  • Book structure: I Ch1–4, II Ch5–9, III Ch10–16, IV Ch17–23, V Ch24–29, VI Ch30–35, VII Ch36–43, VIII Ch44–51, IX Ch52–56, X Ch57–60. Composes Ch52 emitted, 53 conveyed, 54 enforced, 55 booked, 56 ordered, 57 triaged, 58 isolated, 59 published; spine = Zeller/Agans scientific hypothesis–experiment–evidence loop adapted to stochastic systems (hash the unreproducible, 1→3+ trials, “works now” = UNKNOWN). Ch59 probable-cause/contributing-factors split = the model for “bound the proof”. Ch45 diagnostic-case record threads I–X.

PREREQUISITES

Symptom statement + available-freeze with UNKNOWNs + budget (10/60/full) + guide routing + per-drawer records with contracts satisfied.

LOCAL INVARIANTS

  • Budget first with KEPT/SKIPPED stated; route by symptom to drawers never to verdicts; single-variable + declared trials per drawer with OBSERVED/INFERRED labeled; split compounds; bound proof; queue artifacts by ID.

FAILURE MODES (this chapter’s specific ones)

  • Catalog shelving (sixty instruments, zero selection); synthesis invention (new “unified theory” diagnoses); contract-free composition (names without records); budget skipping (full-rigor claims from triage); gap papering; proof inflation (“any AI system, debugged” — refused on this page); spine abandonment; artifact evaporation.

DIAGNOSTIC METHOD (3-6 steps)

  1. Declare budget + KEPT/SKIPPED.
  2. Route symptom through guide to named drawers.
  3. Run the spine per drawer (freeze, hypothesize, intervene, trial, artifact).
  4. Compose honestly (split compounds, name gaps, bound generality, queue IDs).

RESEARCH-DERIVED IDEAS (papers/findings with bounds)

  • Chi, Feltovich & Glaser, Categorization of Physics Problems, Cog Sci 5(2) 1981, pp.121–152 — experts sort by deep structure (governing principle), novices by surface features; unindexed sixty chapters = novice base; drawers re-index by deep structure (what to freeze, what can be established); “symptoms merely suggest.” Bounds: physics/medicine, not debugging; cross-domain transfer poor (hence obsolescence page).
  • Gawande, Checklist Manifesto, 2009 — short written checklists beat expert memory under pressure (experts skip known steps); paper forms = READ-DO, Debugging Checklists = DO-CONFIRM; catch omissions, don’t replace judgment/human verification.
  • Haynes et al., Surgical Safety Checklist, NEJM 360(5) 2009 — WHO 8-hospital study, complications down ~1/3, deaths nearly halved; measured backing for “paper form IS the tool.” Bounds: surgical; transfer to software by argument, not trial.
  • Zeller, Why Programs Fail, 2009 + Agans, Debugging, 2002 (Ch57) — spine = scientific hypothesis–experiment–evidence loop adapted (hash the unreproducible, 1→3+ trials, “works now” = UNKNOWN).

EXPERIMENT / LAB (actual lab, H-structure)

Lab 60 (PROPOSED): workbench audit over three past incidents spanning correctness, cost/latency, agents. H1: one drawer isolates; H2: compound needs split procedure; H3: named gap, no drawer covers. Route via guide, run drawer procedures with ≥3 trials where required, record hit/split/gap with deciding records. “Complete coverage” without the three trials is not completion.

COMPANION TOOL (name + accepts/can-establish/cannot-establish)

Unified Debugging AI Interface — accepts: symptom, freeze + UNKNOWNs, budget, guide routing, per-drawer records + contracts. Can establish: whether this composed investigation is complete under this budget and which drawers contributed what (symptom + budget only). Cannot establish: new diagnoses, cross-symptom generality, future sufficiency, gap freedom; never uses narration, confidence, correlation, single runs, agreement, calm. Paper form closes the book.

PREVENTION ARTIFACT

Composed record (symptom, budget, drawers + H1/H2/H3, contracts y/n, trials, verdict-under-budget, named gaps, generality boundary, artifact/test/clause/monitor IDs) queued or published.

READER OUTCOME (testable phrasing)

Given three past incidents, reader routes each through the selection guide, executes drawer procedures with contract + trial discipline, and fills a coverage table (drawer hit / compound split / named gap with missing instrument specified) — claiming no new diagnosis and no universality.

DEPENDENCIES

All Parts I–X drawers; Ch57–59 budgets; Ch1 spine (first divergence + causal test).

FORWARD BRIDGE

There is no next chapter. Next is the reader’s incident: ten minutes routes, an hour isolates, the full procedure publishes — with the book’s established-vs-provisional split honored.

EVIDENCE / RESEARCH REQUIREMENTS

Reader’s own three-incident audit; constructed “wrong and slow” routing only, no measured runs. Established (in bounds): spine across budgets, per-drawer contracts, routing/bisection/calibration/ledger/publication disciplines. Provisional: cross-system generality, threshold permanence, unknown-class coverage, machine-debugger sufficiency (Part VIII establishes the gates and instruments; whether any assistant meets them is each deployment’s own measurement), method replacing judgment in high-impact calls.

ANTI-CLAIMS / LIMITS

Covers examined symptom + budget + named drawers only; UNKNOWN wherever drawer records absent, trials unrun, or gaps named. No new diagnoses/benchmarks/vendor claims here by contract; “debugging is the scientific method” is framing, not proof.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part X — The Debugging AI Playbook

Sixty chapters, one workbench — and the question of what it proved

Chapter 59 ended with the published packet: one incident fully accounted, generality bounded on the final page. Now the practitioner faces the opposite problem — not one incident with a full procedure but any symptom with sixty chapters behind it. Which instrument, in which order, for this failure? The failure is meta and concrete at once: a teammate asks “the agent is wrong and slow — where do I start?” and the answer cannot be “read the book again.” Sixty chapters that cannot be selected under pressure are sixty chapters shelved.

OBSERVATION: the presenting symptom is compound (wrongness + latency) with no case sentence, no freeze, and no owner; the practitioner holds sixty chapters of instruments with no stated selection order. HYPOTHESIS H1 (selection failure): the knowledge exists but the symptom→instrument mapping is missing. H2 (instrument failure): the mapping exists but the instruments were never contracted for composability. H3 (proof failure): instruments and mapping exist, but nothing states what each combination actually establishes — so selection is faith. INFERENCE: none yet — H1/H2/H3 predict different workbench signatures, and this final chapter retires all three by inventorying the kit, contracting its composition, and bounding its proof.

This chapter’s question: what is the accumulated toolkit as one coherent kit — how is each instrument selected, what does each combination prove, and what does the book leave provisional?

Why “a list of sixty tools” fails first

The obvious move — appending a catalog of chapter summaries — fails because lists do not compose and summaries do not select. Five defects hide behind catalog closings:

  1. Flat inventory. Sixty entries, no parts, no order. The reader with a live incident reads none of them.
  2. Restated diagnoses. The closing invents new failure claims (“a unified theory of AI error”) that no chapter earned. Synthesis must index, not invent — this book claims no new diagnoses here.
  3. Orphaned contracts. Tool names without accepts/performs/can/cannot/next-action. Names without contracts are souvenirs.
  4. Missing selection. No symptom→chapters→tools mapping. The workbench exists; the drawer labels don’t.
  5. Proof inflation. “You can now debug any AI system.” No: the practitioner can execute stated procedures under stated bounds. The difference is the book’s honesty — kept to the final page.

OPINION: a toolkit chapter that teaches new tricks is a betrayal of the sixty that taught the old ones. This chapter indexes, selects, and bounds. Nothing new; everything findable.

The mental model: the workbench with drawer labels — ten parts as ten drawers, each instrument filed with its contract, selected by symptom through the guide below, composed through the book’s one spine: freeze → hypothesize → intervene single-variable → trial → artifact. Every investigation at every budget runs the same spine with the drawers its budget admits (Ch57–59’s KEPT/SKIPPED rules). The kit is coherent because the spine is single; the drawers differ only in depth.

The selection guide exists because of a well-replicated finding about expertise: Chi, Feltovich & Glaser (1981) showed that novices sort problems by surface features (“this one has an inclined plane”) while experts sort by deep structure (“this is a conservation-of-energy problem”). “The agent is wrong and slow” is a surface description; routing it to a budget and then to a drawer by what the evidence could establish is the deep-structure move. The guide is a scaffold for categorizing failures the way an experienced debugger already does — which is why its final rule is “symptoms merely suggest.”

The method: inventory, selection guide, composition contract

The kit in full, by drawer — each entry as contract, not summary:

  1. Part I (Ch1–4): debugging from first principles. The discipline, the first divergence, evidence hygiene, the layer stack. Reach here when the investigation itself is unstructured: no written intent, no pinned repro, explanations ahead of evidence. Outputs: frozen repros, first-divergence statements, convicted layers.
  2. Part II (Ch5–9): deterministic software. Tracebacks, live state, boundaries, contracts, environments. Reach here when the failure may be ordinary code: crashes, wrong values, green suites with boundary holes, works-here-fails-there drift. Outputs: convicted handoffs, boundary tables, handoff assertions, parity records.
  3. Part III (Ch10–16): interactive and numerical AI. Notebooks, data, tensors, training, evaluation. Reach here when state accumulates invisibly or numbers mislead: out-of-order cells, leaky splits, shape/device faults, flat losses, high scores on broken models. Outputs: clean-Run-All records, quarantine artifacts, shape contracts, calibrated evals.
  4. Part IV (Ch17–23): models. Opaque-system stance, pipeline-vs-weights attribution, rendered inputs, truncation ledgers, sampling distributions, triage-only internals, behavioral diffs. Reach here when outputs vary under fixed code or a revision lands: trial distributions, length ledgers, diff gates. Outputs: pinned input records, attribution rows, diff-gated rollouts.
  5. Part V (Ch24–29): AI-assisted development and research. Roles, intent contracts, agent context, designs, research provenance, code-agent trajectories. Reach here when the work product is generated: underspecified asks, unverified citations, stuck agents. Outputs: intent contracts, provenance tables, minimal-trajectory repros.
  6. Part VI (Ch30–35): prompts, retrieval, hallucinations. Versioned prompts, prompt minimization, retrieval stages, retriever-vs-generator splits, claim tables, explanation audits. Reach here when wording, chunks, or citations are suspect: hash diffs, rank-gap records, exoneration ledgers. Outputs: version-blame verdicts, deploy gates, withhold decisions.
  7. Part VII (Ch36–43): agents and trajectories. Trajectory contracts, failure taxonomy, loops, replay/fork, causal replay, trajectory diffs, multi-agent routing. Reach here when agents act: contracted traces, loop breaks, causal replays, boundary maps. Outputs: routed repairs (edge/receiver/coordination).
  8. Part VIII (Ch44–51): building the AI debugger. Oversight gates, crash dumps, machine-checkable invariants, hypothesis enumeration, experiment design, diagnosis verification, benchmark design, debugger self-test. Reach here when delegating diagnosis to models or building debugging machinery — under the contracts, never as narrator-oracle. Outputs: contract-checked machine findings, human-verified; frozen bundles, preregistered ledgers.
  9. Part IX (Ch52–56): production. Per-request records, incident conveyor, guardrails, cost/latency ledgers, live order. Reach here when users or money are involved: emission audits, frozen bundles, calibrated clauses, joint ledgers, cadenced response. Outputs: replayability, durability, enforcement, booked economy, preserved debuggability.
  10. Part X (Ch57–60): playbook budgets. Ten-minute triage, one-hour isolation, full investigation, this workbench. Reach here first — budget selects procedure, procedure selects drawers.
    flowchart TD
    SYM["any symptom (surface description)"] --> BUD{"budget available?"}
    BUD -->|"10 min"| T57["Ch57 triage — route, never verdict"]
    BUD -->|"~1 hour"| T58["Ch58 isolate — bisect + minimize + mini-suite"]
    BUD -->|"days, roles staffed"| T59["Ch59 full investigation — published packet"]
    T57 --> RT["route by deep structure (what the evidence could establish) to the drawer(s)"]
    T58 --> RT
    T59 --> RT
    RT --> DR["10 drawers: Part I discipline ... Part IX production ... Part X budgets"]
    DR --> SP["run the ONE spine per drawer: freeze -> hypothesize -> intervene single-variable -> >=3 trials -> artifact"]
    SP --> CMP{"how does the symptom resolve across drawers?"}
    CMP -->|"compound, spans drawers"| SPLIT["split — one spine per half, joint decision only on joint evidence"]
    CMP -->|"no drawer covers it"| GAP["named gap — workbench output, not shame"]
    CMP -->|"one drawer isolates it"| ART["queue the artifact; bound generality (these drawers, this symptom, this budget)"]
    SPLIT --> ART
  
SELECTION GUIDE (symptom -> chapters -> tools):
wrong output, single case ......... Ch57 triage -> Ch30-35 -> Prompt Program Checklist / rank-gap record
wrong output, repeated ............. Ch58 isolate -> Ch2 bisection -> Ch18-23 attribution -> Incident-to-Regression Replay Builder
agent misbehaved ................... Ch36-43 -> contracted trace / Trajectory Diff / Interaction Map
hallucinated citation .............. Ch34 -> withhold rules + exoneration ledger
money moved / data exposed ......... Ch56 live order -> Ch52 freeze -> Ch53 conveyor -> Ch54 clause
slow or expensive .................. Ch55 ledger + probes -> Cost/Latency Debug Dashboard
"improved" but unmeasured .......... Ch16 eval discipline -> Ch23 diffs -> ablation + benchmark-fidelity record
machine debugger proposed .......... Ch44 gates -> capability checklist, human verification
any incident, full weight .......... Ch59 packet -> Full Incident Investigation Checklist
unsure where to start .............. Ch57 (10 min) -> route, never verdict
RULE: budget first (57/58/59), then drawer. Procedure selects instruments; symptoms merely suggest.

OBSERVATION (constructed illustration, not a measured run): the “wrong and slow” teammate case routes through ten-minute triage (case sentence + freeze + one probe), then splits — wrongness to the Ch30–35 drawer, slowness to the Ch55 ledger — with each half carrying its own predictions and neither half treated as diagnosed. UPDATED BELIEF: H1 retired for this instance (mapping exists — the guide above); H2 retired going forward (composition runs the one spine at every budget); H3 addressed below (proof bounded explicitly). Index, spine, bounds — the workbench stands.

No new diagnosis is claimed in this chapter — by design and by the contract below. Any “unified theory” sentence, any benchmark, any vendor capability, any confidence-scored recommendation appearing here would violate the book’s hard rules at the finish line. The kit composes; it does not coronate.

Example: working the workbench on “wrong and slow”

The practitioner answers the teammate without rereading sixty chapters:

# workbench composition: budget selects procedure, procedure selects drawers
triage = ten_minute(case="agent wrong + slow")  # OBSERVATION + UNKNOWNs + one probe + route
# Split by symptom, one spine each; predictions pre-written per half:
wrongness = isolate(drawer="Ch30-35", spine=[freeze, hypothesize, intervene, trials, artifact])
slowness = attribute(drawer="Ch55", spine=[ledger, baseline, probes, joint_verdict])
# Compose: two records, two owners or two timeboxes; joint decision only on joint evidence.
handoff(records=[wrongness.record, slowness.record], next=[owner_or_queue]*2)

In the constructed case wrongness routes to a prompt-version suspect (hash diff queued for the hour) and slowness routes to a retry-multiplier suspect (ledger queued for probes) — two drawers, one spine, zero verdicts at triage. The licensed claim covers this routing of this symptom — the workbench demonstrates selection, not omniscience.

Lab 60: workbench audit with pre-written coverage predictions (proposed)

PROPOSED, not executed: no author-measured results are reported. The evidence this chapter requires is the reader’s own retrospective audit.

Setup. Take three of your own past incidents (or staged equivalents spanning correctness, cost/latency, and agents). Lay out this chapter’s selection guide beside them. The guide’s routing (correct drawer first try vs. misroute) is the independent variable; incidents and records are controlled.

Task.

  1. Before routing, write H1/H2/H3 with distinct predicted selection signatures: H1: “symptom maps to exactly one drawer whose instrument isolates the divergence”; H2: “symptom spans drawers (compound) and needs the split procedure”; H3: “no drawer covers the symptom — a named gap, honestly recorded.”
  2. Route each incident through the guide; attempt the drawer procedure from its chapter record; run ≥3 trials where the drawer requires trials.
  3. Record coverage: drawer hit, compound split, or named gap — with the record that decided each.
Hypothesis Predicted selection signature FORECAST OBSERVATION (×3 incidents) UPDATED BELIEF
H1 single-drawer one drawer isolates ___ ___ live/exonerated
H2 compound split procedure required ___ ___ live/exonerated
H3 gap no drawer covers ___ ___ live/exonerated

Success criterion. A coverage table over three incidents with routing verdicts, trial-backed drawer outcomes, and any gaps named with the missing instrument specified. A claim of “complete coverage” without the three trials is explicitly not completion.

Companion tool: Unified Debugging AI Interface

What it accepts: any symptom statement, the available-evidence freeze with named UNKNOWNs, the budget (10/60/full), the selection guide’s routing, and the per-drawer records with their contracts. What it performs: it routes budget→procedure→drawers, verifies each drawer’s contract (accepts/performs/can/cannot/next-action) is satisfied before composition, checks every cell ran declared trials with pre-written predictions, confirms OBSERVED-vs-INFERRED labeling throughout, and assembles the composed record with the generality boundary stated — refusing composition where any drawer record is missing or unlabeled. What it can establish: whether the composed investigation is complete under its stated budget and which drawers contributed what — for the examined symptom and budget only. What it cannot establish: new diagnoses, cross-symptom generality, future sufficiency, or freedom from the named gaps. It never treats model narration, confidence, correlation, single runs, agreement, or downstream calm as compositional glue. How its output changes your next action: complete-under-budget routes to the artifact queue or publication; missing-drawer routes to the named chapter procedure; named-gap routes to the backlog as specified instrumentation — each tracked by ID.

Paper form, sufficient for this chapter — and the book:

Symptom ___ | Budget 10/60/full ___ | Routing: drawers ___ (H1/H2/H3: ___)
Records: ___ (contracts satisfied y/n ___, trials ___)
Composed verdict: ___ under budget ___ | Gaps named: ___
GENERALITY: these drawers, this symptom, this budget. Nothing universal.
NEXT: artifact/test/clause/monitor IDs ___

Where a software implementation does not yet exist in the reader’s stack, this record is the tool. The workbench closes; the spine holds.

Research lineage: expertise, checklists, and debugging as method

Why a selection guide and not a summary. Chi, Feltovich & Glaser’s expert–novice studies (Cognitive Science, 1981) established that expertise is largely a matter of how knowledge is organized for retrieval: experts have their domain indexed by the principles that determine the solution, novices by the visible features of the problem. Sixty chapters of instruments with no index is a novice’s knowledge base — everything present, nothing retrievable under load. The drawer/spine structure is a deliberate re-indexing by deep structure (what must be frozen, what can be established) rather than by symptom.

Why the paper forms are the deliverable. Gawande’s The Checklist Manifesto (2009), drawing on the WHO Safe Surgery Saves Lives study (Haynes et al., NEJM 2009 — surgical complications down roughly a third, deaths nearly halved across eight hospitals), makes the case this book has followed for sixty chapters: in complex, high-stakes, time-pressured work, a short written checklist outperforms expert memory, and it does so precisely because the expert under pressure skips steps they know. The per-chapter “paper form, sufficient for this chapter” blocks are READ-DO checklists in Gawande’s sense; the “Debugging Checklist” sections are DO-CONFIRM. The claim is not that checklists replace skill — it is that they catch the predictable omissions skill alone does not.

Debugging as the scientific method. The spine — freeze, hypothesize, intervene on one variable, run trials, record the artifact — is the hypothesis–experiment–evidence loop that Zeller (Why Programs Fail, 2009), Agans (Debugging, 2002), and the delta-debugging line all converge on. This book’s contribution is not the loop; it is adapting each step to a stochastic, opaque, non-stationary system: freeze becomes hashing a non-reproducible input, one trial becomes three-plus, and “it works now” becomes an explicit UNKNOWN.

What the synthesis honestly cannot do. Chi et al. also found that expert categories transfer within a domain and poorly across domains; a debugging toolkit indexed to today’s model families, retrieval stacks, and agent frameworks will need re-indexing as those change. The book states its bounds on the final page for the same reason the NTSB writes a probable cause and a separate contributing-factors list: the honest output names what it established and leaves the rest open.

Bounds: the expertise research is from physics and medicine, not debugging specifically; the checklist evidence is surgical and its transfer to software is by argument, not trial; “debugging is the scientific method” is a framing with wide practitioner support but no controlled proof. The transferable core: index knowledge by deep structure, write the checklist for the step you will skip under pressure, run the scientific loop, and bound the claim.

Reusable procedure: use the workbench on every failure

  1. Budget first — 10, 60, or full; KEPT/SKIPPED stated up front.
  2. Route by symptom — guide rows to drawers, never to verdicts.
  3. Run the spine — freeze, hypothesize, intervene single-variable, trial, artifact — per drawer.
  4. Compose honestly — compound symptoms split; gaps named, never papered.
  5. Bound the proof — generality stated per drawer, per symptom, per budget.

Failure modes

  • Catalog shelving. Sixty instruments, zero selection. The guide exists; use it first.
  • Synthesis invention. New diagnoses smuggled into the closing. Index, don’t invent.
  • Contract-free composition. Drawer names cited without their records. Names are not evidence.
  • Budget skipping. Full-rigor claims from triage time. KEPT/SKIPPED is the price of honesty.
  • Gap papering. Uncovered symptoms forced into drawers. Named gaps are workbench output, not shame.
  • Proof inflation. “Any AI system, debugged.” The book’s final temptation; refused on this page.
  • Spine abandonment. Multi-variable heroics with sixty chapters of warning. The spine is single; stay on it.
  • Artifact evaporation. Brilliant routing, nothing queued. Unqueued workbench sessions never happened.

Limits, per contract: this interface covers the examined symptom under the stated budget with the named drawers; it certifies nothing beyond those bounds and stays UNKNOWN wherever drawer records are absent, trials unrun, or gaps named.

References

  • Michelene Chi, Paul Feltovich, and Robert Glaser. Categorization and Representation of Physics Problems by Experts and Novices. Cognitive Science, 5(2), 1981, pp. 121–152. https://doi.org/10.1207/s15516709cog0502_2
  • Atul Gawande. The Checklist Manifesto: How to Get Things Right. Metropolitan Books, 2009.
  • Alex Haynes et al. A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population. New England Journal of Medicine, 360(5), 2009, pp. 491–499. https://doi.org/10.1056/NEJMsa0810119
  • Andreas Zeller. Why Programs Fail: A Guide to Systematic Debugging (2nd ed.). Morgan Kaufmann, 2009.
  • David Agans. Debugging: The 9 Indispensable Rules. AMACOM, 2002. (Cross-ref Ch 57.)

Debugging Checklist

  • Budget declared (10/60/full) with KEPT vs. SKIPPED stated?
  • Symptom routed through the selection guide to named drawers?
  • Each drawer’s contract (accepts/performs/can/cannot/next-action) satisfied?
  • H1/H2/H3 pre-written per drawer with distinct signatures?
  • Single-variable interventions with ≥3 trials per cell (or budget-labeled fewer)?
  • OBSERVED vs. INFERRED labeled across the composed record?
  • Compound symptoms split (not forced into one drawer)?
  • Gaps named with the missing instrument specified?
  • Generality bounded (these drawers, this symptom, this budget)?
  • Artifacts queued or published by ID?
  • No new diagnoses, benchmarks, vendor claims, narration, scores, single runs, agreement, or calm cited?

What This Chapter Established

  • The accumulated toolkit as one workbench: ten drawers with contracts, the symptom→chapters→tools selection guide, and the single-spine composition rule — demonstrated on the constructed “wrong and slow” routing, no measured runs claimed.
  • What the book established (proved within bounds): the spine (freeze→hypothesize→intervene→trial→artifact) across ten/ sixty/full budgets; per-drawer instruments with contracts; routing, bisection, calibration, ledger, and publication disciplines — each licensed to its examined cases, pins, and trials.
  • What the book left provisional (NOT proved): cross-system generality; threshold permanence across traffic; unknown-class coverage; machine-debugger sufficiency (Part VIII establishes the gates and instruments; whether any assistant meets them is each deployment’s own measurement to make); and any claim that method replaces judgment in high-impact decisions, where human verification remains required.
  • Lab 60 as a proposed workbench audit the reader executes; the Unified Debugging AI Interface contract (accepts/performs/can-establish/cannot-establish/next-action).
  • Position in the arc: Chapter 52 emitted, 53 conveyed, 54 enforced, 55 booked, 56 ordered, 57 triaged, 58 isolated, 59 published — this chapter files every instrument in one workbench with its bounds on the label. Sixty chapters, one spine.
  • Research grounding: expert knowledge is indexed by deep structure not surface features (Chi, Feltovich & Glaser 1981), so the toolkit ships a selection guide; short written checklists beat expert memory under pressure (Gawande; Haynes et al., NEJM 2009), so the paper forms are the deliverable; the spine is the scientific method adapted to a stochastic system.

Next

There is no next chapter. There is the next incident — yours, with its own hashes, its own budget, its own UNKNOWNs. Open the workbench at the budget you have: ten minutes routes, an hour isolates, the full procedure publishes. Freeze before theorizing, move one variable, run the trials, queue the artifact, label what you inferred, and state what you did not prove. The book established the method within its bounds and left the rest provisional on purpose — because debugging AI, done honestly, is the discipline of saying what the evidence shows, naming what it doesn’t, and building the prevention the next incident will need. The method is yours now. Use it under pressure.