When the Frame Is Wrong
Chapter 12 ended with leverage: wrong memory averages 0.042 with two harmful tasks, so a layer that decides what the model may see can cause the harmful action rather than merely failing to prevent it. The next question is when the system should trust its own framing that strongly. Chapter 10 specified the missing piece without building it β a frame that carries how it was established, with conditioning that weakens as the evidence degrades toward query-only retrieval. This chapter builds that piece and measures whether it earns its place.
Establishment is observable; confidence is not
The gate classifies how a frame was established from observable work signals only: corroborated declaration, weak inference, stale basis, conflicting signals, unknown or undeclared provenance β plus an adversarial probe in which a poisoned record rides a declared frame. It never sees hidden labels, and it never trusts scalar confidence or model self-report; the outside literature gives that refusal its backing, and the Research foundations section names it. Each class maps to a policy action: hard framing, soft framing without exclusion, query-only fallback, broadened retrieval, requesting more evidence, or abstention with a forced-abstain path where action would be destructive.
Twenty fixtures in twelve families (AβL) encode the hazard list: declared decisions, weak inferences, shifted objectives, conflicting evidence, unknown projects, poisoned listings, unanswerable unknowns, and an irrelevant task that needs no memory at all. Eleven form the development split, nine the held-out evaluation split, both frozen by digest. Adaptive-retrieval thresholds were inspected on development data and deliberately retained unchanged β the corpus margins top out at 0.016, so any dominance threshold below 0.20 would tune to noise β and one proposed mapping revision (conflicting frames broaden instead of falling back) was replayed on development data and rejected: a single case with nil task effect cannot carry a mapping change, though the observation behind it, that fallback deleted a facade the broader context kept, is recorded as a finding about fallback risk.
The ladder
Every scenario runs an F-ladder under one reader at temperature zero. No-memory (F0) anchors attribution. Fallback (F1) is the byte-equivalent query-only path. Correct-hard (F2) and wrong-hard (F3) carry the true and deliberately wrong frames. Soft (F4) and broadened (FB) vary exclusion and ranking. The gate’s choice composes F6 from the already-read condition. An adaptive baseline (F7) fills from query-only ranking by pre-registered branching rules. A frame-sceptic baseline (F8) abstains wherever establishment does not justify action. An oracle (FO) with hidden truth takes the best permitted action, abstaining exactly where abstention is acceptable.
Contexts reuse frozen Chapter 12 rows wherever the constructed string is character-identical, hash-gated per reader. That last qualifier was earned the hard way: the first second-reader attempt reused llama outcomes under ministral contexts because the lookup keyed context without reader identity β 41 contaminated rows, caught by audit, fixed by keying reuse on task, hash and reader, and re-run clean with a regression test pinning the rule. The contaminated run is discarded, not repaired.
Book result (frozen runs
ch13-dev-v1,ch13-eval-v1onllama3.1:8b,ch13-eval-v1-ministralonministral-3:8b, grader v2, thresholds frozen, mapping v1). The gate matches its pre-registered class on all 9 evaluation scenarios for both readers, with zero false hard-framings, zero harmful F6 tasks, and no gate breaches. Benefit retention holds on its single eligible pair with gap 0.0. The evaluation contrasts, llama primary:
scenario gateβF6 F6 F2 F1 F0 FO
A arch HARD 1.0 1.0 1.0 0.75 1.0
B fix SOFT 1.0 0.833 0.833 0.333 1.0
C shift QUERY_ONLY 1.0 0.75 1.0 0.75 1.0
D cite QUERY_ONLY 0.25 0.0 0.25 0.0 0.25
E ship QUERY_ONLY 0.167 β 0.167 0.0 0.167
F prose QUERY_ONLY 0.0 β 0.0 0.0 0.0
G fix REQUEST 0.333 β 0.833 0.333 0.833
K fix REQUEST 0.333 β 0.833 0.333 0.333
L irrel SOFT 0.0 β 0.0 1.0 β
F6 matches the oracle on 7 of 8 oracle-eligible scenarios; the exception is G, an unnecessary abstention on an answerable task. Soft framing beats hard framing on B (1.0 against 0.833, oracle-matched), and query-only fallback beats the true but stale frame on C (1.0 against 0.75, oracle-matched). The second reader reproduces the gate profile exactly β 9/9 classes, 7/8 oracle matches with the same single G exception, no breaches β at different magnitudes that the chapter reports separately and never averages.
The headline contrast is a measured nil. Correct-hard against wrong-hard separates on no family under the pre-registered 0.25 bar on either reader; ministral gaps are exactly 0.0 everywhere. This is not a failed manipulation: verified bundle rebuilds show the wrong-frame contexts are near-disjoint from the correct ones (unit overlap 0.10, 0.23, 0.00 and 0.21 across the four pairs), yet behaviour does not move β and on the stale-shift task the wrong arch frame scores 1.0 against the true publication frame’s 0.75. Against the outcomes declared before the run this is a Type D result: partial credit, with real wins, real costs, and a dead headline contrast.
What the nil means
Work-type frame errors do not propagate to behaviour here because query support absorbs them: the load-bearing units admit any work type, so replacing seventeen of seventeen context units leaves the reader’s actions unchanged. The frame hazards that bite are elsewhere. Staleness bites (C: the true frame underperforms no-memory’s fallback). Poisoning bites on development data (the fallback listing deletes the contracted facade where the gated context keeps it). Over-confidence bites (G, L). A wrong label on the frame is survivable; a wrong belief about what the frame is worth acting on is not. That reframes the chapter’s own thesis the way the strong-reader programme anticipated: the problem was never basic classification accuracy but tail risk β graceful fallback, safe degradation, conflict, staleness, abstention, and reader suggestibility.
The costs are findings too
Requesting more evidence costs half a task on G for both readers (0.333 against an answering oracle at 0.833β1.0), inside the pre-registered tolerance of one but real. Framing an unframeable task costs everything on L for the primary reader: every framed condition refuses while no-memory echoes correctly (0.0 against 1.0) β memory displacing obedience to the present task, the same intrusion Chapter 12 recorded on its echo control. The second reader shows no such cost (1.0 under frames), so over-conditioning is reader-dependent, not architectural: the policy induces refusal in one reader and compliance in the other, and the chapter cannot say which it will induce in a third. Broader retrieval reaches 0.5 on E where gate and oracle alike sit at 0.167 β upside outside the oracle mapping, a direction for the mapping’s next version, not a breach. F is a published floor (0.0 everywhere including oracle). Blanket scepticism loses everywhere it can be compared: F8 matches or trails no-memory on all nine tasks.
Transfer
Five time-locked repository questions at the frozen commit, scored separately from fixtures, return a null with a diagnosis: memory-supplied and memory-free answers score 0.0β0.5 throughout, and the single half-credit (UNKNOWN establishment, destructive action retained) comes from the no-memory condition. The reader does not reproduce the establishment taxonomy unaided because the taxonomy never appears in the evidence β the transfer task as designed measures exact-value emission of labels the reader was never given, so its zeros are uninterpretable as capability evidence and are reported as a misspecified probe, not a capability verdict.
What remains unsolved. The strong-reader lane stays open: the frame policy is frozen and reader-independent by design, so a capable third reader can be compared without retuning, and the handoff records what it should run. The adaptive baseline’s skip branch never fired live and remains unit-test-only coverage. Poison-blocking rests on development data alone with no evaluation probe. The broaden-upside on unknown frames wants a mapping version the frozen evaluation could not adopt. Over-conditioning needs a mechanism account of why one reader refuses where another complies. And the prose task that floors every condition including the oracle is kept as a published instrument failure.
Research foundations
The outside literature converges on distrusting the signals a frame gate must not use, and on the policy surface this chapter builds. Moskvoretskii and colleagues (AdaRAGUE, ACL 2025) compare dozens of adaptive-retrieval and uncertainty methods and treat model self-confidence as suspect. Min and colleagues (QuCo-RAG, ACL Findings 2026) declare model-internal confidence unreliable and trigger retrieval from corpus statistics instead. Huang and colleagues (UncertaiNLP 2025) frame abstention as a control action under cost asymmetry, the closest precedent to the REQUEST path. Yan and colleagues (CRAG, 2024, preprint) build a lightweight retrieval evaluator with triggered corrective actions, the closest prior art to a frame-quality gate. Jeong and colleagues (Adaptive-RAG, NAACL 2024) route among retrieval strategies by question complexity, and Jiang and colleagues (FLARE, EMNLP 2023) separate when to retrieve from what to retrieve β the two precedents for the adaptive baseline the policy must beat rather than merely differ from. Srivastava and He (MemoryGraft, 2025, preprint) show poisoned experience persisting through retrieval, motivating the poisoning probe.
References
- Viktor Moskvoretskii and colleagues, Adaptive Retrieval without Self-Knowledge? Bringing Uncertainty Back Home (ACL 2025).
- Dehai Min and colleagues, QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation (ACL Findings 2026).
- Zhiqi Huang and colleagues, Confidence-Based Response Abstinence: Improving LLM Trustworthiness via Activation-Based Uncertainty Estimation (UncertaiNLP 2025).
- Shi-Qi Yan and colleagues, Corrective Retrieval Augmented Generation (2024, preprint).
- Soyeong Jeong and colleagues, Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity (NAACL 2024).
- Zhengbao Jiang and colleagues, Active Retrieval Augmented Generation (EMNLP 2023).
- Saksham Sahai Srivastava and Haoyu He, MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval (2025, preprint).