Chapter 17 of 60

Debugging What You Cannot See

Concepts

CHAPTER 17 β€” DEBUGGING WHAT YOU CANNOT SEE

PART IV β€” Debugging Models

PURPOSE

Opens Part IV on the certified-but-failing intent slice (wrong split-shipment refund, no leak/metric/seed explanation): installs the opacity + stochasticity stance β€” interiors uninspectable, boundaries fully pinnable, trial series (β‰₯5) as the permanent unit of evidence β€” plus the Part IV vantage map and the accreting frozen-bundle record.

CENTRAL QUESTION

When the mechanism is not directly observable, what discipline replaces “open it and look”?

UNIQUE CLAIM

Uninspectable does not mean undiagnosable: diagnosis moves to a behavioral contract around the glass β€” exact input bytes, parameters, outputs pinned in a hashed frozen bundle, perturbed one crossing at a time with pre-written numeric FORECASTs over β‰₯5 trials β€” because generated explanations are sampled post-hoc outputs with no causal access (Ch3’s Turpin/Lanham constraint as founding rule), transparency is absent by field property (Lipton), post-hoc explainers are unfaithful (Rudin) and gameable via off-distribution probes (Slack/LIME/SHAP), so the robust response is in-distribution boundary perturbation counted over series; Part IV’s chapters are then one object (observable boundary behavior around the same opaque computation) from six vantage points (attributionβ†’renderβ†’windowβ†’distributionβ†’signalsβ†’diffs), with the bundle accreting (render + token ids in Ch19, length ledger in Ch20, sampling + series in Ch21, pinned suite in Ch23).

DEBUGGING OBJECT

Evidence as boundary behavior β€” frozen bundles (revision + input bytes + params + seed + outputs, hashed) and fixture counts (2/12 β†’ 11/12 on the split-shipment set); H1 retrieval vs H2 instruction vs H3 weights separated only by boundary interventions.

CONCEPTS INTRODUCED

Boundary-first diagnosis (define-as-count β†’ freeze β†’ rival hypotheses β†’ single discriminating intervention β†’ pin); frozen bundle as the structured hashed diagnostic case; transparency-vs-post-hoc-explanation distinction (Lipton frame); in-distribution vs off-distribution perturbation (why this probe differs from LIME/SHAP); verdict-creep quarantine (boundary convictions stay scoped to fixture+revision).

CONCEPTS DEVELOPED / REUSED

Self-report-is-not-trace from Ch3 (promoted to Part IV founding constraint with Turpin/Lanham cross-ref); CheckList behavioral stance from Ch16 (generalized to boundary perturbation); β‰₯3-run habit from Ch12/Ch15/Ch16 (hardened to β‰₯5-trial floor for stochastic probes); bundleβ†’suite lineage foreshadowed (prevention representation for later parts).

PREREQUISITES

Ch3 (evidence hygiene, load-bearing test), Ch16 (behavioral testing, instrument discipline), Ch12 (repeats), Ch1 (rival hypotheses + FORECASTs).

LOCAL INVARIANTS

State behavior as fixture count + criterion; freeze the bundle before any swap; every hypothesis names its killing boundary intervention (“confused model” is not a hypothesis); one variable per probe, β‰₯5 trials (single improved answer = UNKNOWN); non-matching patterns = UNKNOWN with next probe named; pin survivor as fixture + bundle, file the dead hypothesis.

FAILURE MODES

Explanation-as-trace (self-report quoted as cause); single-run conviction (one retry closing a stochastic case); multi-variable repair (prompt+context+temperature moved together, nothing learned); boundary amnesia (unrecorded revision/inputs making all comparisons UNKNOWN); interior surrender (“can’t see inside, can’t debug”); fixture-of-one (no count/rate/FORECAST); verdict creep (retrieval-miss retold as “can’t reason about exceptions”).

DIAGNOSTIC METHOD

  1. Count the behavior (fixture, criterion, current score). 2. Freeze bundle + hashes. 3. Write H1/H2 with mutually exclusive numeric FORECASTs. 4. Move one boundary variable, β‰₯5 trials, compare. 5. Pin survivor as regression; route (retrieval/instruction β†’ Ch18–19; all-boundary-survived β†’ interior formally suspect, boundary certified).

RESEARCH-DERIVED IDEAS

Lipton CACM 2018 (Mythos: interpretability underspecified; transparency vs post-hoc explanation β€” Part IV assumes no transparency); Rudin Nature MI 2019 (black-box post-hoc explanations frequently unfaithful; high-stakes β†’ inherently interpretable models; stuck-with-black-box β†’ explanation layer is no diagnostic); Slack et al. AIES 2020 (adversarial models with clean LIME/SHAP via off-distribution probing β€” motivates in-distribution changes + series counting); Rai et al. 2024 practical mechanistic-interpretability review (SAEs scaled to frontier models; transcoders/crosscoders trace features β€” but the SAE unit is seed-unstable (different init β†’ substantially different features), not all latents interpretable, and circuit hypotheses still need boundary/causal validation; research-grade, model-specific, not a debugging instrument for a hosted endpoint β€” the stance holds with a footnote).

EXPERIMENT / LAB

Lab 17 (PROPOSED): repeatable model failure (or injected missing-doc / contradicting-instruction), frozen bundle, 8–12 case fixture, H1/H2 exclusive FORECASTs (e.g. repaired-context β‰₯10/12 vs ≀4/12), baseline Γ—5 + probe Γ—5. H-structure: independent var = one boundary change; controls = revision/seed/everything else. Success = pattern-matching table + pinned regression bundle; changed-answer-without-table is not completion.

COMPANION TOOL

Black-Box Hypothesis Tester β€” accepts: behavior definition + criterion + frozen bundle + H1/H2 FORECASTs + trial series. Can-establish: which boundary hypothesis survives this fixture under this revision β€” nothing about the interior. Cannot-establish: weight mechanisms, fixture-external generality, future revisions; never model-explanations/single-answers/paraphrase-agreement as evidence.

PREVENTION ARTIFACT

Regression fixture + pinned bundle hashes + filed dead hypotheses with exonerating rows (Ch8 contract applied to a dark interior).

READER OUTCOME

Reader can run a scoped black-box conviction without ever citing the model’s story β€” testable via Lab 17’s hypothesis table.

DEPENDENCIES

Ch3, Ch16, Ch12, Ch1.

FORWARD BRIDGE

Ch18 “Is the Model Actually the Problem?” β€” inherits the continent problem: “the boundary” is still prompt/retrieval/context/params/weights, and the glass hides all five equally.

EVIDENCE / RESEARCH REQUIREMENTS

2/12β†’11/12 repaired-context illustration constructed; opacity argued as field property not tooling gap; no interior/weights/representation claims licensed here.

ANTI-CLAIMS / LIMITS

One table convicts one boundary cause under one revision+fixture; explains no weights, certifies no model, survives no revision bump un-rerun; UNKNOWN wherever single-trial or multi-variable.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part IV β€” Debugging Models

The glass wall

Chapter 16 ended with an honest instrument and an unanswered question. The repaired evaluation reports, repeatably: the model fails the intent slice β€” wrong refund answer on split shipments, confident, cited, and wrong. No leak explains it. No metric swap rescues it. No seed relitigates it.

So the engineer does the natural thing. She opens the model and looks inside β€” and finds there is no inside to open. No stack frame at the wrong line. No variable holding the wrong value. No refund_id key to watch go missing at a handoff. There is a weights file measured in gigabytes, an API endpoint, and a text answer. The failure is behind glass: observable only through inputs and outputs, never by inspection.

Two properties of this system compound, and Part IV is organized around both. The first is opacity: the mechanism cannot be watched, so diagnosis moves to the boundary. The second is stochasticity: each output is one draw, so the unit of evidence is no longer a value but a distribution of behavior under a pinned input β€” the β‰₯3-run habit Chapters 12, 15, and 16 built, now permanent and primary. You debug less of the mechanism with more of the statistics. Everything in this chapter about pinning the boundary assumes you are pinning it across a trial series, not a single call.

OBSERVATION: identical frozen input reproduces the wrong answer across reruns at temperature 0 (MEASUREMENT, n=5, same model revision); no intermediate state is exposable by any reader-available probe. HYPOTHESIS H1 (wrong retrieval): the context fed to the model lacks the split-shipment policy. HYPOTHESIS H2 (wrong instruction): the system prompt tells the model to prefer the general policy over the exception. HYPOTHESIS H3 (wrong weights): the model cannot apply the exception even when shown it. INFERENCE: none yet β€” all three predict the identical wrong paragraph. Only interventions on the boundary (not introspection of the interior) separate them.

This chapter’s question: when the mechanism is not directly observable, what discipline replaces “open it and look”?

The shape of Part IV

Part IV debugs one thing from several angles: the boundary between an opaque model and the observable system around it. The mechanism never comes into view, so each chapter picks a different observable site on that boundary and works it with controlled experiments, moving from the model’s input side to its output side and then across revisions:

  • Chapter 18 β€” which layer. Reproduce the case, then attribute the failure to a subsystem: pipeline, parameters, weights, or an intent gap.
  • Chapter 19 β€” the rendered input. Inspect the exact bytes and token ids the model received, not the prompt you wrote.
  • Chapter 20 β€” the context-window boundary. Account for every token the window kept and every one it silently cut.
  • Chapter 21 β€” the behavioral distribution. Treat the spread of outputs under a pinned input as the object, and learn how to characterize and compare it.
  • Chapter 22 β€” internal diagnostic signals. Use logprobs, entropy, and attention to decide where to probe next β€” never as an explanation of why.
  • Chapter 23 β€” the good-run / bad-run diff. Turn the measured difference between two revisions (or two runs) into a gated regression asset.
    flowchart LR
    PR["prompt + system instruction"] --> G
    RET["retrieved documents"] --> G
    ASM["assembled context + token ids"] --> G
    PAR["sampling parameters"] --> G
    G["opaque weights β€” the interior never testifies"] --> DIST["behavioral distribution over a trial series"]
    G --> SIG["logprobs / entropy / attention signals"]
    DIST -.->|"diff two revisions"| REG["pinned regression suite"]
  

The nominal object changes every chapter. The real object does not: it is always observable boundary behavior around the same opaque computation. Chapters that look like a change of subject are a change of vantage point, chosen deliberately. This chapter fixes the stance β€” why an uninspectable, stochastic system still yields to controlled boundary experiments; Chapter 21 turns the stochastic half of that stance into method.

Watch one thing accumulate as the Part proceeds. This chapter introduces a frozen bundle: model revision, input bytes, parameters, seed, and outputs, all hashed. Chapter 19 adds the render and token ids; Chapter 20 adds the length ledger; Chapter 21 adds the sampling configuration and trial series; Chapter 23 promotes a set of bundles into a pinned suite two revisions can be run against identically. A reproducible diagnostic case stops being an ad-hoc script and becomes a structured, hashed record. That representation matters later.

Why asking the model fails first

The obvious move β€” asking the model why it answered wrong β€” fails because a generated explanation is another output, not a trace. It is sampled text shaped to sound plausible, produced after the answer, with no causal access to the computation that produced the answer. Treating it as a stack trace is the founding error of Part IV.

Chapter 3 already established this for chain-of-thought specifically β€” Turpin and colleagues’ biased-CoT result, Lanham and colleagues’ load-bearing test. Part IV takes the same rule as its founding constraint: the interior does not testify.

Three opacity traps follow from that error:

  1. Post-hoc storytelling. “I prioritized the general policy because…” reads like a confession and functions like fiction. Change the question order, and the story changes while the weights do not.
  2. Single-run inference. One failure, one theory, one fix. Stochastic systems punish this faster than deterministic ones ever did β€” the next sample contradicts the theory for free.
  3. Multi-variable thrash. New prompt and new documents and higher temperature in one retry. If the answer improves, nothing was learned; if it worsens, nothing was learned louder.
  4. Anecdote archiving. Screenshots of one failure in a ticket thread, no bundle, no seed. Six weeks later nobody can re-run anything β€” the incident is a story, not a fixture.

OPINION: uninspectable does not mean undiagnosable. It means the diagnosis lives at the boundary β€” in controlled inputs and measured outputs β€” instead of in the interior. The team that accepts this early spends its instrumentation budget on boundary fidelity (bundles, hashes, trial counts) instead of on interior theater.

The mental model: a behavioral contract around an uninspectable interior. You cannot watch the gears, so you pin what crosses the glass in both directions β€” exact input bytes, exact parameters, exact outputs β€” and you debug by perturbing one crossing at a time and predicting the movement before you look. The interior never testifies; the boundary always does.

The method: boundary-first diagnosis

The five-step discipline, unchanged from the book loop but hardened for opacity:

  1. Define the failing behavior with a measurable success criterion. Not “bad refund answer” but “on the 12-case split-shipment fixture, the answer cites policy section 4.2-exception and computes the correct amount; currently 2/12.”
  2. Freeze the boundary bundle. Model revision + exact input bytes + sampling parameters + output text + seed, all hashed. Without this, later “improvements” are incomparable β€” Chapter 12’s reproduction rule, now mandatory because there is nothing else to hold still.
  3. Generate competing hypotheses from preserved evidence only. Every hypothesis must name a boundary intervention that would separate it. “The model is confused” is not a hypothesis; “replacing the retrieved section changes the citation” is.
  4. Run one discriminating intervention with a pre-written prediction. One variable moves; the FORECAST says which hypothesis lives or dies on each possible outcome. Because sampling varies, stochastic probes run β‰₯5 trials; a single improved answer is UNKNOWN, not success.
  5. Verify, document, convert to prevention. The surviving hypothesis earns a regression fixture plus the boundary bundle β€” Chapter 8’s contract applied to a system whose interior stays dark.
# black-box hypothesis test (interior never inspected; boundary fully pinned)
bundle = freeze(model_rev="rev-2026-08-14", input_bytes=rendered, params={"temperature": 0, "seed": 7})
baseline = run(bundle, trials=5)  # MEASUREMENT: 2/12 cite the exception
# H1 probe: same model, repaired retrieval (one variable moves)
probe_h1 = run(bundle.with_context(sections=["4.2-exception"]), trials=5)
# FORECAST: H1: probe_h1 >= 10/12; H2: unchanged until system prompt moves; H3: unchanged under any context
print("baseline:", baseline, "probe_h1:", probe_h1)
# All other patterns -> UNKNOWN; next probe named, not concluded.

OBSERVATION (constructed illustration, not a measured run): baseline 2/12; repaired-context probe 11/12 across 5 trials at temperature 0. UPDATED BELIEF: H1 supported for this fixture and revision; H2/H3 suspended, not deleted β€” a retrieval defect and a weights defect can coexist. INFERENCE: the fix belongs outside the weights (retrieval repair + fixture), and no interior story was needed to place it there.

Note what the illustration does and does not claim. It claims that on this fixture, under this revision, the repaired-context probe moves the count as H1 forecast β€” a boundary fact, checkable by re-running. It does not claim the model “understands” the exception, that retrieval is the only defect in the system, or that the next revision will behave the same. Each of those is a fresh HYPOTHESIS awaiting its own probe series, and the regression bundle exists so the series can run without re-arguing the setup.

Research lineage: why the boundary is the right stance

The opacity this chapter accepts is not a temporary tooling gap; it is a well-argued property of the field.

“Interpretable” is underspecified. Lipton catalogued the many incompatible things researchers mean by model interpretability and separated two: transparency (you can actually trace the mechanism) and post-hoc explanation (a separate artifact produced after the fact that describes the model) (Lipton, 2018). Part IV assumes no transparency and treats every post-hoc explanation β€” the model’s own, a saliency map, a surrogate β€” as an artifact to be tested, not trusted.

Post-hoc explanations of black boxes can mislead. Rudin argues that explanations approximating a black box are frequently unfaithful to it, and that for high-stakes decisions the right move is an inherently interpretable model rather than a black box plus an explanation (Rudin, 2019). When you are stuck with the black box, her critique is a warning: the explanation layer is not a diagnostic instrument.

Even the dedicated explanation tools are fragile. Slack and colleagues built models that behave badly on real inputs but present innocuous LIME and SHAP explanations, exploiting the fact that those methods probe the model with off-distribution perturbations (Slack et al., 2020). The constructive response, and this book’s, is behavioral: perturb the real boundary with in-distribution changes, predict the movement, and count outcomes over a trial series β€” the CheckList stance from Chapter 16, generalized.

The interior is not entirely dark β€” but it is not your instrument either. Mechanistic interpretability has made real progress: sparse autoencoders have been scaled to frontier-size models to pull out human-readable features, and follow-on techniques (transcoders, cross-layer decompositions) trace how those features connect (Rai et al., 2024). But three facts keep this out of the debugging loop for now. It is model-specific and research-grade β€” there is no drop-in interior probe for an arbitrary hosted endpoint. Its core unit is unstable: sparse autoencoders trained on the same model and data with different random seeds learn substantially different feature sets, so “the feature that fired” is partly an artifact of the decomposition. And a feature or circuit hypothesis is still a hypothesis β€” it earns the word cause only under a causal intervention, which has its own pitfalls (Chapter 22). The practitioners who can partly read the interior still validate every claim at the boundary. So the stance holds, with a footnote: the glass is getting less opaque in the lab, and none of that changes what you do on Tuesday.

Lab 17: one intervention, two hypotheses, pre-written predictions

PROPOSED, not executed: no author-measured results are reported. The evidence this chapter requires is the reader’s own hypothesis table.

Setup. Take one repeatable model failure (or inject one: remove the decisive document from the context, or append a contradicting instruction to the system prompt). Freeze the boundary bundle β€” revision, input bytes, params, seed. Fix the success criterion as a count on a small fixture (8–12 cases), not a feeling.

Task.

  1. Write H1/H2 with mutually exclusive numeric FORECASTs before intervening (e.g., “H1 retrieval: repaired-context run scores β‰₯10/12; H2 instruction: repaired-context run stays ≀4/12”).
  2. Independent variable: exactly one boundary change (context or instruction or params). Controlled variables: everything else pinned, including model revision and seed.
  3. Execute baseline (β‰₯5 trials) and probe (β‰₯5 trials). Record OBSERVATION (outputs verbatim or hashed) and UPDATED BELIEF per hypothesis. Any pattern outside both FORECASTs is UNKNOWN with the next probe named.
  4. Convert the surviving hypothesis into a regression fixture + pinned bundle; file the dead hypothesis with its exonerating row instead of deleting it.
Row Condition FORECAST OBSERVATION UPDATED BELIEF
baseline frozen bundle Γ—5 H1/H2 agree: ≀4/12 ___ none yet
probe one boundary change Γ—5 H1: β‰₯10/12; H2: ≀4/12 ___ H1 live / H2 live / UNKNOWN

Success criterion. A completed table matching one pre-written pattern plus the pinned regression bundle. A changed answer without the table is explicitly not completion β€” motion is not diagnosis.

Companion tool: Black-Box Hypothesis Tester

What it accepts: the failing-behavior definition + success criterion, the frozen boundary bundle (revision, input bytes, params, seed hashes), two competing hypotheses with numeric FORECASTs, and the trial series (β‰₯5 runs per condition). What it performs: it enforces order (freeze β†’ FORECAST β†’ intervene β†’ compare), checks observed counts against FORECASTs, refuses a verdict on single runs or multi-variable probes, and stamps the surviving hypothesis with bundle hashes. What it can establish: which boundary hypothesis survives this fixture under this revision β€” and nothing about the interior mechanism. What it cannot establish: why the weights behave so, whether the verdict generalizes beyond the fixture, or future validity under a new revision. It never treats a model-generated explanation, a single improved answer, or agreement across paraphrases as evidence. How its output changes your next action: a surviving retrieval/instruction hypothesis routes to Chapters 18–19 (attribution, then input inspection); a hypothesis that survives all boundary probes routes forward with the interior formally suspect and the boundary certified.

Paper form, sufficient for this chapter:

BEHAVIOR: ___ (fixture ___ cases; criterion ___)   BUNDLE: rev ___ | input hash ___ | params ___ | seed ___
H1: ___ FORECAST: ___   H2: ___ FORECAST: ___
BASELINE Γ—5: ___   PROBE Γ—5 (one var: ___): ___
CONVICTION: H1 / H2 / UNKNOWN (pattern-match line: ___)
REGRESSION: fixture ___ + bundle hashes ___

Where a software implementation does not yet exist in the reader’s stack, this record is the tool. The boundary discipline precedes any automation.

Reusable procedure: every opaque failure gets this first

  1. State the behavior as a count β€” fixture, criterion, current score.
  2. Freeze the bundle β€” revision, input bytes, params, seed, output hashes.
  3. Write rival hypotheses with FORECASTs β€” each naming its killing intervention.
  4. Move one boundary variable, β‰₯5 trials β€” compare against FORECASTs.
  5. Pin the survivor as regression β€” fixture + bundle, dead hypotheses filed.

Failure modes

  • Explanation-as-trace. Quoting the model’s self-report as the cause. It is a second symptom wearing a cause’s clothes β€” file it as output, never as mechanism.
  • Single-run conviction. One retry passes, case closed. Nondeterministic systems require distributions; one draw proves nothing except that one draw happened.
  • Multi-variable repair. Fixing prompt, context, and temperature together and declaring victory. The symptom moved; the defect kept its address.
  • Boundary amnesia. Debugging without the frozen bundle β€” revision unrecorded, input reconstructed from memory. Every later comparison is then UNKNOWN by construction.
  • Interior surrender. “We can’t see inside, so we can’t debug.” Opacity removes one instrument (inspection), not the discipline (controlled boundary experiment).
  • Fixture-of-one. A single failing example treated as the behavior definition. One case cannot carry a count, a rate, or a FORECAST β€” build the 8–12 case fixture before theorizing.
  • Verdict creep. A boundary conviction (“retrieval missed 4.2 here”) retold by Friday as a mechanism claim (“the model can’t reason about exceptions”). Convictions stay scoped to fixture and revision or they become folklore.

Limits, per contract: one hypothesis table convicts one boundary cause under one revision and fixture; it does not explain the weights, does not certify the model, and does not survive a revision bump without re-running. UNKNOWN wherever trials are single or variables moved together.

References

  • Zachary C. Lipton. The Mythos of Model Interpretability. Communications of the ACM 61(10), 2018, pp. 36–43 (first version arXiv:1606.03490, 2016). https://doi.org/10.1145/3233231
  • Cynthia Rudin. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nature Machine Intelligence 1, 2019, pp. 206–215. https://doi.org/10.1038/s42256-019-0048-x
  • Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. Fooling LIME and SHAP: Adversarial Attacks on Post Hoc Explanation Methods. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2020, pp. 180–186. https://doi.org/10.1145/3375627.3375830
  • Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, and Ziyu Yao. A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models. arXiv:2407.02646, 2024. https://arxiv.org/abs/2407.02646

Debugging Checklist

  • Failing behavior stated as fixture count + success criterion?
  • Boundary bundle frozen (revision, input bytes, params, seed, output hashes)?
  • H1/H2 with mutually exclusive numeric FORECASTs written before intervening?
  • Exactly one boundary variable moved, β‰₯5 trials per condition?
  • OBSERVATION recorded verbatim/hashed; UPDATED BELIEF per hypothesis?
  • UNKNOWN declared where patterns match neither FORECAST?
  • Survivor pinned as regression fixture + bundle?

What This Chapter Established

  • The opacity stance: interiors uninspectable, boundaries fully pinnable; diagnosis lives in controlled input/output perturbations, never in model-generated explanations.
  • The boundary-first discipline (define β†’ freeze β†’ rival hypotheses β†’ single discriminating intervention β†’ pin), demonstrated on the split-shipment case as H1-supported β€” constructed illustration, no measured runs claimed.
  • Lab 17 as a proposed hypothesis record the reader executes; the Black-Box Hypothesis Tester contract (accepts/performs/can-establish/cannot-establish/next-action).
  • What was NOT proved: any claim about weights, representations, or mechanisms β€” this chapter convicts boundary causes only.
  • Research grounding: opacity is a well-argued property, not a tooling gap (Lipton on transparency vs. post-hoc explanation); post-hoc explanations of black boxes can be unfaithful (Rudin) and even LIME/SHAP can be adversarially fooled via off-distribution probing (Slack et al.). Mechanistic interpretability is opening parts of the interior in the lab (Rai et al.) but its unit is seed-unstable and every finding is still boundary-validated, so the practitioner stance is unchanged. The robust response is in-distribution boundary perturbation with pre-written forecasts.
  • Part IV’s opening map: this chapter sets the stance; Chapters 18–23 apply it layer by layer (attribution, input, truncation, sampling, signals, diffs).

Next

The stance is set β€” but “the boundary” is still a continent. The wrong answer could come from the prompt, the retrieved documents, the assembled context, the sampling parameters, or the weights themselves, and the glass hides all of them equally. The next chapter narrows the suspect list: is the model actually the problem, or is everything around it?