Chapter 13 of 17

When the Frame Is Wrong

Concepts

Chapter 13 β€” When the Frame Is Wrong

Source: 13-chapter.md

What this chapter is really about

Underneath the F-ladder, this chapter is about trusting framing by evidence rather than by default. A frame that arrives declared is acted on hard; one that arrives inferred is acted on softly; one that arrives stale, conflicting, or unknown is weakened toward query-only retrieval, broadened search, or an explicit request for more evidence. Its deepest move is splitting two quantities the earlier chapters fused: whether the frame label is correct against whether acting on the frame is safe. The run shows they come apart β€” near-disjoint wrong-frame contexts leave behaviour unchanged, while stale and over-confident frames cost real task score. The hidden thesis: establishment evidence, not frame correctness, is the load-bearing quantity.

Current thesis

Explicit claims

  • Frame establishment (corroborated/weak/stale/conflicting/unknown, from observable signals only, never confidence or self-report) decides policy action: HARD/SOFT/QUERY_ONLY/BROADEN/REQUEST/ABSTAIN.
  • F2-vs-F3 positive control with a 0.25 bar gates interpretation: work-type frame errors do not move behaviour on these tasks (near-disjoint bundles, gaps ≀0.125 llama, exactly 0.0 ministral).
  • Gate policy earns partial credit (Type D): soft win on B, stale-avoidance win on C (llama), oracle-matched retention on A/D/K (+E/F), measured costs on G (unnecessary abstention, both readers) and L (over-conditioning, llama-only).
  • F6 matches the oracle on 7/8 eligible scenarios for both readers; readers reported separately, never averaged.
  • Transfer is a reported null with a misspecification diagnosis (taxonomy absent from evidence), not capability evidence.

Implied claims

  • Observable work signals suffice to establish frames well enough to act on: 9/9 gate-class match, zero false hard-framings, on both readers.
  • Query support absorbs work-type error: load-bearing units admit any work type, which is the mechanism behind the F2/F3 nil.
  • Over-conditioning is reader-dependent (refusal in llama, compliance in ministral), so framing policy cannot promise refusal behaviour across readers.
  • Frozen thresholds and mapping transfer across readers without retuning (same policy, both readers, no per-reader tuning by design).

Not yet established

  • That the establishment taxonomy transfers to unaided reader classification (transfer probe misspecified; still open).
  • That SKIP-branch adaptivity ever fires live (unit-test-only coverage).
  • That poison-blocking generalises beyond the single development probe (no eval J scenario).
  • Why one reader refuses framed irrelevant tasks while another complies (mechanism unknown).
  • Whether broaden-on-unknown beats query-only systematically (E upside is single-task).

What the chapter already gives us

  • The establishment gate as a versioned object. Classes, signals, actions, and the v1 mapping frozen before evaluation with a recorded backtest rejection β€” policy as an auditable artefact, not a posture.
  • The positive-control discipline applied inward. The F2/F3 gate does to this chapter what Chapter 12’s BW did to the pipeline: it forbids the headline the authors wanted (wrong frames hurt) and forces the chapter onto what the numbers support.
  • The reuse-contamination catch. 41 reader-mixed rows found by audit, fixed by reader-keyed reuse, re-run clean, regression-pinned β€” frozen-artifact discipline surviving contact with a second reader.
  • Reader-dependence mapped, not averaged. Same gate profile on both readers with different magnitudes and two reader-local effects (L cost, E upside) β€” the instrument measures memory–reader interaction, continuing Chapter 12’s finding.
  • The tuning-freeze record. Provisional thresholds retained with stated reason (dev margins top out at 0.016), backtest rejected with stated reason (n=1, nil effect) β€” pre-registration with visible teeth.

Where the current treatment stops

  • The trace format’s consumer is still unspecified (fixture scorers, human debuggers, self-monitor each need different content) β€” carried over unresolved from the previous treatment.
  • The policy↔maintenance cycle (stale state degrades policy which degrades maintenance) is named in the old concepts but unmapped by the run; staleness appears only as a fixture class, not as dynamics.
  • Trust/deny signals for adversarial history stop at one development probe; usefulness-policy and safety-policy remain unintegrated.
  • The oracle mapping itself may be incomplete: broaden beats oracle-permitted actions on E, so FO is a ceiling over permitted actions, not over achievable behaviour.
  • as_of discipline (decision vs production-state timestamps) never enters the gate; temporal staleness is detected by signal conflict, not by timestamp comparison.

The deeper territory

  • Establishment vs correctness as separate axes. The run’s central surprise deserves formalising: under what retrieval conditions does label-correctness stop predicting behaviour? Query-support breadth is the current mechanism; its boundary (narrow-support tasks where F2/F3 should separate) is untested.
  • Refusal as a reader property. L shows framing induces refusal in one reader and not another. Is refusal propensity measurable per reader in advance (F0-vs-framed gap on irrelevant probes as a reader fingerprint)? If so, REQUEST/ABSTAIN policies should condition on the reader, breaking the reader-independence the frozen policy assumes.
  • The oracle-mapping gap. E suggests permitted-action oracles understate achievable behaviour. A fuller oracle (best over all conditions, not best permitted) would change FO-match accounting β€” at the cost of oracle interpretability.
  • Policy tuning splits as standard. The dev/eval fixture split plus frozen thresholds worked here; generalising it to every hand-set number in the book (the old concepts’ manifest-audit idea) remains future methodology work.

Concepts worth developing

Establishment evidence as the load-bearing quantity

Idea. Replace frame-correctness with establishment-quality as the variable policies condition on: declared/inferred/stale/conflicting/unknown as versioned, observable, auditable state carried with every frame.

Why it matters. The run shows correctness without establishment-awareness costs (C: true-but-stale frame loses to fallback) and establishment-awareness without correctness wins (C again, from the other side).

Connection to the current chapter. Generalises the F0–F5 gate into frame metadata carried by Chapters 10–11 outputs.

Broader implication. Any system that conditions strongly on inferred structure needs an establishment channel alongside the structure β€” frames, beliefs, extractions alike.

What remains unresolved. Timestamp-based staleness vs signal-conflict staleness; who supplies as_of when tasks do not.

Reader fingerprints for refusal policy

Idea. Probe each reader with framed irrelevant tasks (L-family) before deploying REQUEST/ABSTAIN-heavy policies; condition abstention aggressiveness on the measured refusal propensity.

Why it matters. L costs 1.0 task score on llama and 0.0 on ministral β€” the same policy is safe and destructive depending on reader.

Connection to the current chapter. Turns the reader-dependence finding into a deployment input.

Broader implication. Breaks strict reader-independence of memory policy; the frozen policy may need a reader-calibration annex.

What remains unresolved. Stability of refusal propensity across prompts/tasks; whether calibration transfers across policy versions.

Permitted-action vs achievable-behaviour oracles

Idea. Report both oracles: best permitted action (current FO, interpretable) and best condition overall (E-style upside detector).

Why it matters. E’s broaden upside is invisible to permitted-oracle accounting; the mapping’s next version needs the signal.

Connection to the current chapter. Extends the FO semantics already recorded.

Broader implication. Oracle design is itself a policy choice that shapes verdicts; Type-D accounting should show both.

What remains unresolved. Whether achievable-oracles invite post-hoc mapping fitting (the Β§68 held-out-set rule would govern adoption).

Important distinctions

  • Establishment quality versus frame correctness (the chapter’s central split).
  • Permitted-action oracle versus best-condition oracle (E exposes the gap).
  • Reader-independent policy versus reader-calibrated abstention (L forces the question).
  • Tuning freeze with stated reason versus tuning to noise (thresholds) and underpowered promotion versus discipline (backtest).
  • Contaminated reuse versus segregated reuse (the 41-row catch).
  • Misspecified probe versus capability verdict (transfer null).
  • Insensitive family versus ineffective policy (the positive-control cut, inherited and applied).

What mechanism would make this work?

Deterministic gate over observable signals with versioned mapping; F-ladder with frozen reuse keyed by task, hash and reader; dev/eval fixture split with frozen thresholds; F2/F3 positive control with pre-registered bar; oracle with abstention-acceptable semantics; adaptive baseline with recorded branches; frame-sceptic and fallback baselines; time-locked transfer with verbatim verification. Missing: live SKIP coverage, eval poison probe, timestamp-based staleness, reader calibration, trust/deny integration.

Connections to the rest of the book

  • Consumes Chapter 10’s frames and Chapter 12’s instrument/grader/tasks; builds the abstention path Chapter 10 specified.
  • Hands Chapter 14 tight-budget behaviour (left to this chapter) back measured; hands Chapters 15–17 a versioned policy object plus the establishment-vocabulary.
  • Continues Chapter 11’s operating-point thinking (thresholds as policy parameters) and Chapters 5/7 auditability (trace per decision).
  • The 41-row reuse catch hardens the frozen-run contract Chapter 2 demands.
  • Chapter 17’s insertion point (canonical run now exists, Type D) is updated by this chapter.

Beyond the current book

  • Adaptive retrieval as policy surface (FLARE, Adaptive-RAG, CRAG, AdaRAGUE, D2-RAG, QuCo-RAG β€” chapter refs).
  • Abstention as control action under cost asymmetry (confidence-abstinence line).
  • Memory poisoning via retrieved experience (MemoryGraft β€” probe motivation).
  • Reader suggestibility and calibration (extends Chapters 12–13 reader-dependence into deployment practice).

Verified entries already cover the policy surface and the distrusted signals; the bandit/RL formalisms stay deferred until a learned block exists. Open literature passes: timestamp-based staleness detection in RAG pipelines; reader refusal-propensity measurement; permitted-vs-achievable oracle design in policy evaluation.

Possible future claims

Already supportable

  • Establishment-gated framing with frozen mapping matches an oracle on 7/8 eligible cases across two readers.
  • Work-type frame errors do not move behaviour on these tasks despite near-disjoint contexts.
  • Requesting evidence costs measurably on answerable unknowns; framing unframeable tasks induces reader-dependent refusal.

Plausible but needs development

  • Broaden-on-unknown beats query-only systematically (single-task upside).
  • Reader fingerprints predict abstention-policy safety per deployment.
  • Establishment metadata carried with frames generalises beyond fixtures.

Speculative

  • Timestamp-based staleness detection subsumes signal-conflict classes.
  • Permitted-oracle gaps systematically locate mapping improvements.
  • The F2/F3 nil extends to narrow-support tasks (untested; mechanism predicts it should not).

Claims worth attacking

  • “The gate earns its place.” Counter: on ministral, gate wins shrink to retention (B/C at ceiling/parity) β€” the policy’s value may concentrate in weaker readers, and the chapter’s wins may be reader-specific rescue rather than general mechanism.
  • “The F2/F3 nil is robustness.” Counter: it may be task-selection artefact β€” all four pairs use broad-support tasks where query support rescues; narrow-support tasks could separate sharply, and none were run.
  • “Transfer null is misspecification.” Counter: convenient β€” a failed transfer is relabelled uninformative. The defence is the mechanism (taxonomy absent from evidence), but an independent replication with taxonomy present would be needed to close the charge.

Tensions and counterarguments

  • Reader-independence (frozen policy, no per-reader tuning) vs reader-calibrated abstention (L demands calibration).
  • Pre-registered bars vs post-hoc mechanism stories (C wrong-beats-true was not predicted; reported, not promoted).
  • Auditable frozen reuse vs multi-reader reuse complexity (the contamination catch raises the cost of every future second reader).
  • Partial credit (Type D) vs maximal claims: the chapter’s wins tempt a safety story the costs forbid.

Examples and thought experiments

  • The stale-shift task: true publication frame scores 0.75, wrong arch frame 1.0, query-only 1.0. Which frame is “correct” β€” and does the question matter once establishment says stale?
  • The unnecessary abstention: gate says REQUEST on an answerable fix task (0.333 vs 0.833). Price the abstention: how many G-cases per prevented B6 before REQUEST earns its keep?
  • The reader fingerprint: llama refuses framed irrelevant tasks, ministral complies. Before deploying this policy on a third reader, what single probe would you run?
  • The broaden upside: FB reaches 0.5 where oracle-permitted actions sit at 0.167. Is the oracle wrong, or is the mapping incomplete?

Potential demonstrations or experiments

Done: F-ladder both readers, positive control with bundle-overlap diagnostics, adaptive baseline with recorded branches, backtest rejection, transfer with verbatim lock. Proposed, none run: eval poison probe; SKIP-branch live coverage; taxonomy-present transfer replication; narrow-support F2/F3 pairs; reader refusal fingerprinting.

Research questions this chapter creates

  • Under what retrieval conditions does frame-label correctness stop predicting behaviour (boundary of the nil)?
  • Is refusal propensity a stable per-reader property measurable in advance?
  • Do permitted-action oracles systematically hide mapping improvements across policy evaluations?
  • What timestamp discipline would make staleness detection independent of signal conflict?

Architectural implications

  • Frames should carry establishment metadata as versioned state, not bare labels.
  • Reuse registries must key on reader identity wherever readers share contexts.
  • Abstention-heavy policies need reader calibration before deployment claims.
  • Mapping revisions require fresh held-out sets (Β§68 honoured in the backtest rejection).
  • Transfer probes must include the taxonomy they grade in the evidence.

How would we know this works?

The chapter works if the gate matches pre-registered classes without false hard-framings, retains benefit where frames are sound, wins where frames are stale or weak, costs visibly where it abstains or over-frames, and reports a dead headline contrast with the mechanism (bundle overlap) rather than burying it β€” with all of it reproduced in direction by a second reader. It fails usefully by showing exactly which of those clauses break: G breaks retention-by-abstention, L breaks reader-independence, transfer breaks taxonomy portability.

The chapter at its highest level

The ideal version would teach: establishment over correctness as the policy variable; positive controls that forbid wanted headlines; reuse discipline that survives second readers; reader-dependence reported rather than averaged; costs published beside wins with the same precision. The current version measures all five; the ideal version would also close the open lanes (poison eval, SKIP coverage, taxonomy transfer, refusal mechanism).

Discussion

Start here

  • Near-disjoint wrong-frame contexts leave behaviour unchanged. Is this robustness of behaviour β€” or insensitivity of our tasks? What narrow-support task would decide?
  • The gate abstains unnecessarily on G for both readers. Price it: how many prevented harms fund one wasted abstention?
  • Llama refuses framed irrelevant tasks; ministral complies. Is reader-independence of the frozen policy still defensible?

Push the idea further

  • If establishment metadata rode every frame from Chapters 10–11 onward, which later mechanisms (consolidation, forgetting, procedures) would condition on it first?
  • Reuse keyed by reader saved this chapter’s second reader. What else in the frozen-run contract silently assumes one reader?
  • The transfer probe failed by omitting its taxonomy from the evidence. How many of the book’s nulls are probe artefacts wearing capability verdicts?

Decisions we need to make

  • Whether reader calibration (refusal fingerprinting) enters policy deployment or stays a research question.
  • Whether permitted-action oracles get an achievable-behaviour companion in future verdict accounting.
  • Whether the eval poison probe and SKIP live coverage gate Chapter 13’s final acceptance or ship as documented gaps.

Claims worth attacking

  • “Establishment evidence suffices for safe framing.” Counter: 9/9 class match on fixtures is thin; real staleness is gradual, not the crisp STALE fixtures β€” the gate may pass fixtures and fail the world.
  • “Type D is the honest verdict.” Counter: with B/C wins llama-only and costs G/L/E spread across readers, D may average over a policy whose value concentrates in weak readers β€” report the interaction, not the grade.
  • “The transfer null is uninformative.” Counter: it may instead show establishment taxonomies do not survive contact with readers at all β€” a deeper failure than misspecification.

New ideas worth exploring

  • Establishment metadata as first-class frame state across Chapters 10–17.
  • Reader refusal fingerprinting as deployment input.
  • Dual-oracle verdict accounting (permitted vs achievable).

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Chapter 12 ended with leverage: wrong memory averages 0.042 with two harmful tasks, so a layer that decides what the model may see can cause the harmful action rather than merely failing to prevent it. The next question is when the system should trust its own framing that strongly. Chapter 10 specified the missing piece without building it β€” a frame that carries how it was established, with conditioning that weakens as the evidence degrades toward query-only retrieval. This chapter builds that piece and measures whether it earns its place.

Establishment is observable; confidence is not

The gate classifies how a frame was established from observable work signals only: corroborated declaration, weak inference, stale basis, conflicting signals, unknown or undeclared provenance β€” plus an adversarial probe in which a poisoned record rides a declared frame. It never sees hidden labels, and it never trusts scalar confidence or model self-report; the outside literature gives that refusal its backing, and the Research foundations section names it. Each class maps to a policy action: hard framing, soft framing without exclusion, query-only fallback, broadened retrieval, requesting more evidence, or abstention with a forced-abstain path where action would be destructive.

Twenty fixtures in twelve families (A–L) encode the hazard list: declared decisions, weak inferences, shifted objectives, conflicting evidence, unknown projects, poisoned listings, unanswerable unknowns, and an irrelevant task that needs no memory at all. Eleven form the development split, nine the held-out evaluation split, both frozen by digest. Adaptive-retrieval thresholds were inspected on development data and deliberately retained unchanged β€” the corpus margins top out at 0.016, so any dominance threshold below 0.20 would tune to noise β€” and one proposed mapping revision (conflicting frames broaden instead of falling back) was replayed on development data and rejected: a single case with nil task effect cannot carry a mapping change, though the observation behind it, that fallback deleted a facade the broader context kept, is recorded as a finding about fallback risk.

The ladder

Every scenario runs an F-ladder under one reader at temperature zero. No-memory (F0) anchors attribution. Fallback (F1) is the byte-equivalent query-only path. Correct-hard (F2) and wrong-hard (F3) carry the true and deliberately wrong frames. Soft (F4) and broadened (FB) vary exclusion and ranking. The gate’s choice composes F6 from the already-read condition. An adaptive baseline (F7) fills from query-only ranking by pre-registered branching rules. A frame-sceptic baseline (F8) abstains wherever establishment does not justify action. An oracle (FO) with hidden truth takes the best permitted action, abstaining exactly where abstention is acceptable.

Contexts reuse frozen Chapter 12 rows wherever the constructed string is character-identical, hash-gated per reader. That last qualifier was earned the hard way: the first second-reader attempt reused llama outcomes under ministral contexts because the lookup keyed context without reader identity β€” 41 contaminated rows, caught by audit, fixed by keying reuse on task, hash and reader, and re-run clean with a regression test pinning the rule. The contaminated run is discarded, not repaired.

Book result (frozen runs ch13-dev-v1, ch13-eval-v1 on llama3.1:8b, ch13-eval-v1-ministral on ministral-3:8b, grader v2, thresholds frozen, mapping v1). The gate matches its pre-registered class on all 9 evaluation scenarios for both readers, with zero false hard-framings, zero harmful F6 tasks, and no gate breaches. Benefit retention holds on its single eligible pair with gap 0.0. The evaluation contrasts, llama primary:

scenario   gate→F6        F6    F2    F1    F0    FO
A arch     HARD           1.0   1.0   1.0   0.75  1.0
B fix      SOFT           1.0   0.833 0.833 0.333 1.0
C shift    QUERY_ONLY     1.0   0.75  1.0   0.75  1.0
D cite     QUERY_ONLY     0.25  0.0   0.25  0.0   0.25
E ship     QUERY_ONLY     0.167 β€”     0.167 0.0   0.167
F prose    QUERY_ONLY     0.0   β€”     0.0   0.0   0.0
G fix      REQUEST        0.333 β€”     0.833 0.333 0.833
K fix      REQUEST        0.333 β€”     0.833 0.333 0.333
L irrel    SOFT           0.0   β€”     0.0   1.0   β€”

F6 matches the oracle on 7 of 8 oracle-eligible scenarios; the exception is G, an unnecessary abstention on an answerable task. Soft framing beats hard framing on B (1.0 against 0.833, oracle-matched), and query-only fallback beats the true but stale frame on C (1.0 against 0.75, oracle-matched). The second reader reproduces the gate profile exactly β€” 9/9 classes, 7/8 oracle matches with the same single G exception, no breaches β€” at different magnitudes that the chapter reports separately and never averages.

The headline contrast is a measured nil. Correct-hard against wrong-hard separates on no family under the pre-registered 0.25 bar on either reader; ministral gaps are exactly 0.0 everywhere. This is not a failed manipulation: verified bundle rebuilds show the wrong-frame contexts are near-disjoint from the correct ones (unit overlap 0.10, 0.23, 0.00 and 0.21 across the four pairs), yet behaviour does not move β€” and on the stale-shift task the wrong arch frame scores 1.0 against the true publication frame’s 0.75. Against the outcomes declared before the run this is a Type D result: partial credit, with real wins, real costs, and a dead headline contrast.

What the nil means

Work-type frame errors do not propagate to behaviour here because query support absorbs them: the load-bearing units admit any work type, so replacing seventeen of seventeen context units leaves the reader’s actions unchanged. The frame hazards that bite are elsewhere. Staleness bites (C: the true frame underperforms no-memory’s fallback). Poisoning bites on development data (the fallback listing deletes the contracted facade where the gated context keeps it). Over-confidence bites (G, L). A wrong label on the frame is survivable; a wrong belief about what the frame is worth acting on is not. That reframes the chapter’s own thesis the way the strong-reader programme anticipated: the problem was never basic classification accuracy but tail risk β€” graceful fallback, safe degradation, conflict, staleness, abstention, and reader suggestibility.

The costs are findings too

Requesting more evidence costs half a task on G for both readers (0.333 against an answering oracle at 0.833–1.0), inside the pre-registered tolerance of one but real. Framing an unframeable task costs everything on L for the primary reader: every framed condition refuses while no-memory echoes correctly (0.0 against 1.0) β€” memory displacing obedience to the present task, the same intrusion Chapter 12 recorded on its echo control. The second reader shows no such cost (1.0 under frames), so over-conditioning is reader-dependent, not architectural: the policy induces refusal in one reader and compliance in the other, and the chapter cannot say which it will induce in a third. Broader retrieval reaches 0.5 on E where gate and oracle alike sit at 0.167 β€” upside outside the oracle mapping, a direction for the mapping’s next version, not a breach. F is a published floor (0.0 everywhere including oracle). Blanket scepticism loses everywhere it can be compared: F8 matches or trails no-memory on all nine tasks.

Transfer

Five time-locked repository questions at the frozen commit, scored separately from fixtures, return a null with a diagnosis: memory-supplied and memory-free answers score 0.0–0.5 throughout, and the single half-credit (UNKNOWN establishment, destructive action retained) comes from the no-memory condition. The reader does not reproduce the establishment taxonomy unaided because the taxonomy never appears in the evidence β€” the transfer task as designed measures exact-value emission of labels the reader was never given, so its zeros are uninterpretable as capability evidence and are reported as a misspecified probe, not a capability verdict.

What remains unsolved. The strong-reader lane stays open: the frame policy is frozen and reader-independent by design, so a capable third reader can be compared without retuning, and the handoff records what it should run. The adaptive baseline’s skip branch never fired live and remains unit-test-only coverage. Poison-blocking rests on development data alone with no evaluation probe. The broaden-upside on unknown frames wants a mapping version the frozen evaluation could not adopt. Over-conditioning needs a mechanism account of why one reader refuses where another complies. And the prose task that floors every condition including the oracle is kept as a published instrument failure.

Research foundations

The outside literature converges on distrusting the signals a frame gate must not use, and on the policy surface this chapter builds. Moskvoretskii and colleagues (AdaRAGUE, ACL 2025) compare dozens of adaptive-retrieval and uncertainty methods and treat model self-confidence as suspect. Min and colleagues (QuCo-RAG, ACL Findings 2026) declare model-internal confidence unreliable and trigger retrieval from corpus statistics instead. Huang and colleagues (UncertaiNLP 2025) frame abstention as a control action under cost asymmetry, the closest precedent to the REQUEST path. Yan and colleagues (CRAG, 2024, preprint) build a lightweight retrieval evaluator with triggered corrective actions, the closest prior art to a frame-quality gate. Jeong and colleagues (Adaptive-RAG, NAACL 2024) route among retrieval strategies by question complexity, and Jiang and colleagues (FLARE, EMNLP 2023) separate when to retrieve from what to retrieve β€” the two precedents for the adaptive baseline the policy must beat rather than merely differ from. Srivastava and He (MemoryGraft, 2025, preprint) show poisoned experience persisting through retrieval, motivating the poisoning probe.

References

  • Viktor Moskvoretskii and colleagues, Adaptive Retrieval without Self-Knowledge? Bringing Uncertainty Back Home (ACL 2025).
  • Dehai Min and colleagues, QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation (ACL Findings 2026).
  • Zhiqi Huang and colleagues, Confidence-Based Response Abstinence: Improving LLM Trustworthiness via Activation-Based Uncertainty Estimation (UncertaiNLP 2025).
  • Shi-Qi Yan and colleagues, Corrective Retrieval Augmented Generation (2024, preprint).
  • Soyeong Jeong and colleagues, Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity (NAACL 2024).
  • Zhengbao Jiang and colleagues, Active Retrieval Augmented Generation (EMNLP 2023).
  • Saksham Sahai Srivastava and Haoyu He, MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval (2025, preprint).