Appendix C — Experiments and Reproducibility
What was tested for this edition, what each result supports, and how to reproduce it.
All protocols, fixtures, seeds, runners, and frozen results live in the book’s research companion (research/experiments/). Each family has a spec frozen before its run, a frozen result file, and explicit allowed/forbidden claims. All model-backed runs used small local models, identified below by their recorded names (parameter counts from the local model registry). Summary:
- E02 — Retrieved vs exposed. No model call. Under a tight context budget (60 units; an 80-unit trial had proved too loose, and the tighter protocol was frozen before its run), retrieval-order and gated assembly exposed the decisive span in all runs; distractor-first assembly in none. Retrieval order was fixed by construction. Supports: ordering policy affects exposure under these fixtures and budgets. Does not support: influence on model output, or how often real systems fail this way.
- E04 — Structured handoffs. Schema handoffs retained 70/70 required fields as a fixture property (free text 55/70) and through two blinded model reproducers: qwen2.5:0.5b (free text 60/70) and phi4-mini, 3.8B (free text 56/70); no unsupported additions. Supports: the fixed shape carried tested fields more reliably. Does not support: better task completion.
- E05 — Judge reliability. Candidates from qwen2.5:0.5b. The qwen2.5:0.5b self and same-model judges agreed with a mechanical reference on 0/25 fields (verdict-format failure); a phi4-mini (3.8B) judge agreed on 17/25 with 7 false accepts vs 1 false reject. Supports: judge configurations differ sharply; report false accepts and rejects separately. Does not support: any claim about model size or self- versus other-judging (both changed together), general leniency, or a single accuracy number.
- E07 — Memory admission. qwen2.5:0.5b. Gated retrieval answered 5/8 vs raw 3/8 with zero good notes rejected; one flagged conflicting claim was still selected. The gate read the fixture’s ground-truth note tags, and two answers were scored wrong by the substring checker in both arms. Supports: admission decisions changed outcomes on these fixtures (a mechanism demonstration). Does not support: gates in general, a realistic admission benchmark, or flagging as sufficient.
- E08 — Equal-budget single vs multi. qwen2.5:0.5b, then gemma3 (4.3B). Exact matching scored 0/6 in every arm, twice; a semantic scorer of the same setup scored 5/6 in every arm. The semantic rules were written after the failing outputs were read, frozen before the run, and not validated against human labels. Supports: evaluator design moved the headline (0/6 vs 5/6 from the checker alone); no architecture difference was distinguished on these fixtures. Does not support: any ranking of architectures, or the semantic checker as ground truth.
- E11 — Conformance vs preservation. Separation exists mechanically (6/6 lossy specs conform while failing intent). A blinded single-grader pilot scored 8/8 matches to the construction, but the grader was the author, who also designed the items, so this is not independent human replication. One stronger model channel (gemma3, 4.3B) reproduced the direction on 7/8 items, while a weaker channel (phi4-mini, 3.8B) affirmed everything. Supports: conformance and preservation are separable judgments; grader capability matters. Does not support: prevalence, human agreement, or independent human validation (multi-grader replication remains open).
Reproduce
Deterministic families rerun exactly (run_e02.py, run_e02v2.py, run_e04.py, run_e11.py). Model-backed families record model identity, temperature, seeds, token counts, budgets, and — from v2 on — raw outputs. Frozen files are never overwritten; new protocol versions get new files.