Chapter 27 of 30

Replicate Before You Believe

Concepts

CHAPTER 27 โ€” Replicate Before You Believe

STATUS

Editorial enhancement pass 2026-09-14. Signal framing corrected to the preregistration’s own; “rules out no effect” overclaim replaced by a labeled post-hoc permutation check; token premium decomposed; manipulated-variable provenance examined in the frozen export; exporter defect fixed in CodeAI (uncommitted, regression test, full suite 451 passed). No frozen evidence altered, no model call.

CENTRAL QUESTION

Does the promising prompt intervention survive a matched replication?

THESIS

Signal != result != default. P3 (normal x12 vs counterfactual x12, same model/tasks) tied on the rule’s metrics (9/12, 44/144, identical solve sets, zero rescues) at +45.8% tokens; Rule 2 keeps normal. Counterfactual changed where success appeared without improving the promotion metric; twelve draws per task cannot distinguish that movement from chance (post-hoc shuffle p ~ 0.31).

THE SIGNAL P3 TESTED (corrected)

  • Prereg framing: within P2, counterfactual 16/36 candidates, 9/12 tasks vs normal STANCE (3 draws) 11/36, 7/12 โ€” matched at 3 draws, one 12-task draw, per-stance attribution report-only (P2 rows lack stance field).
  • Ch26’s warned reading: counterfactual 9/12 from 3 draws vs normal ARM 10/12 from 12 โ€” attributable, unmatched.
  • P3 fixes both: 12 v 12, attributable arms, one intended variable.

LOAD-BEARING CLAIMS

  1. P3C/P3CF 9/12, 44/144, tokens 28,325 / 41,303 (+45.8%), input 20,244 / 26,580, output 8,081 / 14,723, p50 3,146 / 3,701.5 ms; identical sets; zero rescues; 288 calls, 286 checks. [measured]
  2. Rule 2 (approx equal -> noise, keep normal) and Rule 6 (stop framing attacks on retry-once-accepted, size-format). “Token premium alone rejects” is the report’s judgment; prereg rules contain no token clause. [reported: prereg, results]
  3. Premium mechanism: 12,978 extra = 6,336 input (44 x 144) + 6,642 output; CF outputs longer on all 12 tasks. [measured: rows]
  4. Frozen export arm configs identical except name (no prompt_suffixes / stance_labels); suffix text absent from export; ledger experiment.created stores suffixes but export_experiment dropped them. Constant +44 input tokens per task per call is the rows’ only trace. [measured + source]
  5. Per-task table recomputed, matches report (3->9, 10->6, 10->8, 6->8, 6->5, 1->2, 3->2, 4->3, 1->1, 0,0,0); sum |diff| = 18. [measured]
  6. Post-hoc (not preregistered) hypergeometric shuffle null on per-task exchangeability: p(sum|diff| >= 18) ~ 0.31 (200k trials, seed 1); unsound-cache Fisher exact ~0.039 uncorrected, one of eight differing tasks. Report reads it as interaction noise. [measured recomputation]
  7. Two EXECUTION_ERROR candidates without check IDs (P3C stable-priority, P3CF strip-query); prompt_version seeded-code-v1 quirk. [measured]
  8. Nosek (prediction vs postdiction), Dodge (budget behind comparisons); mappings ours. HN control anecdote compressed; evidences nothing.

CODE CHANGE (this pass, CodeAI working tree, uncommitted)

  • src/codeai/analysis.py export_experiment: arms now export prompt_suffixes and stance_labels (additive; format string unchanged).
  • tests/test_stances.py::test_export_carries_the_prompt_variable_of_each_arm: builds P3C/P3CF-shaped arms, runs CF, asserts export distinguishes arms and otherwise-identical configs.
  • Old behavior: export arms {name, models, samples, temperature, seed, prompt_version}. New: + prompt_suffixes, stance_labels; exported call rows gain prompt_variant joined from call.requested (collect_experiment). Remaining limit: frozen P2/P3 exports unchanged (historical).
  • Validation: tests/test_stances.py + test_experiments.py 26 passed; full suite 451 passed.

EDITORIAL CORRECTIONS

  • Opening signal framing (was only “9 from 3 vs 10 from 12”).
  • “rules out the strongest dismissal … that the prompt had no effect” removed; replaced by user-specified accurate lesson + shuffle check.
  • Token-premium rejection relabeled as report judgment, not a frozen rule.
  • Teaching toy replaced by executable shuffle check (constructed analysis, real rows).
  • Anecdote section compressed into one paragraph.

EVIDENCE

  • p-series/p3/p3-export.json (experiment.arms, calls input/output tokens, candidates outcomes); p-series-analysis TABLES.md P3 rows; P3-prereg.md; P3-results.md; stances.py; experiments.py ArmDef; analysis.py export_experiment.

CONTROLS / LIMITATIONS

One local model, 12 tasks, one wording, within-corpus; post-hoc test not preregistered; variable attested indirectly in frozen export; two execution errors; token threshold absent from rules.

DEPENDENCIES

Ch26 (P2 subgroup, report-only stance attribution, discovery != promotion), Ch25 (paired rows), Ch24 (denominators, oracle != selector), Ch18 (reported vs measured), Ch13 (token semantics).

FORWARD BRIDGE

Part 5 spent calls freely; Ch28 asks whether spending a call should itself be governed (operation before model).

OPEN ITEMS

  • Export per-call prompt_variant.
  • Ch26 check: its sentence “never designed to estimate counterfactual-at-3 against normal-at-3” conflicts with the portfolio containing normal x3 (report 7/12); fix in the Ch26 pass.
  • CodeAI change awaits user review/commit.

DIAGRAM (2026-09-14)

Added signal to freeze/preregister to new-draws to rule-comparison to promote/reject flowchart. Matches the chapter’s signal/result/default distinction; postdiction never flows backward.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part 5 โ€” More Intelligence Is Not Automatically Better

The challenger gets its matched fight

Chapter 26 ended with a subgroup that looked like a better prompt and a refusal to promote it. This chapter runs the test that refusal called for, and states the answer before interpreting it: at twelve draws per task on both sides, counterfactual wording covered 9 of 12 tasks, matching normal wording, passed 44 of 144 candidates, also matching, while using 46% more tokens. Normal stays the default.

Does the promising prompt intervention survive a matched replication?

What the signal actually was

It is worth being exact about what P3 was built to test, because the signal had two readings and both were weaker than they looked.

The P3 preregistration states the signal in its own words: within one P2 run, counterfactual framing reached 16 of 36 candidates and 9 of 12 tasks, while normal framing in the same portfolio reached 11 of 36 and 7 of 12. That comparison is matched โ€” three draws per task each โ€” but it comes from one twelve-task sample, and its per-stance attribution is not a recorded field in the frozen P2 rows; Chapter 26 recovered it from input-token signatures. The other reading, the one Chapter 26 warned against, sets counterfactual’s 9 tasks from three draws beside the normal arm’s 10 tasks from twelve: attributable, but unmatched. 1 2

P3 removes both weaknesses at once. Each arm gets twelve draws per task, on the same twelve tasks, with the same model, starting state, verifier, sealing, and budget accounting, and each arm is its own attributable arm. The only intended difference is the prompt suffix. The preregistration asks whether the prompt distribution is better, not whether a portfolio is. 1

Signal, result, default

The distinction is a signal โ‰  a result โ‰  a default. A signal suggests a question. It becomes a result only when a frozen question with a frozen decision rule meets new draws. A result changes a default only if the rule says so.

The discipline as a one-way flow โ€” postdiction never flows backward into prediction:

    flowchart TD
    SG["exploratory signal<br/><i>found in frozen rows: postdiction</i>"] --> FR["freeze hypothesis + rule<br/><i>preregister before new draws</i>"]
    FR --> ND["new draws<br/><i>same tasks, one variable changed</i>"]
    ND --> CMP{"comparison<br/>against the frozen rule?"}
    CMP -->|"rule met"| PRO["promote<br/><i>result moves the default</i>"]
    CMP -->|"rule not met"| REJ["reject<br/><i>signal stays a signal</i>"]
  

Preregistration literature supplies the vocabulary for that cycle. Analyses chosen before seeing outcomes test hypotheses; analyses shaped by the data generate them. Treating the second as the first, helped along by ordinary hindsight bias, is how attractive subgroups become false findings (Nosek et al., 2018). The counterfactual subgroup was discovered in P2’s rows, so it was postdiction; P3 is the prediction built from it. That mapping is the book’s own.

The second discipline is reporting the budget behind a comparison. When conclusions shift with the computation spent, a test score alone cannot say which method is better, and the answer can depend on the budget chosen (Dodge et al., 2019). That paper is about model comparisons, not prompt wordings, but the principle transfers: coverage parity at unequal token spend is not parity of methods, so the premium belongs in the headline.

The experiment as code

In CodeAI, a matched prompt replication is two arms that differ in one field. Expressed in the current ArmDef API, the P3 design is:

cf = stance_suffix(COUNTERFACTUAL)
arms = (
    ArmDef(name="P3C",  models=("qwen-local",), samples=12),
    ArmDef(name="P3CF", models=("qwen-local",), samples=12,
           prompt_suffixes=(cf,) * 12,
           stance_labels=(COUNTERFACTUAL,) * 12),
)

The suffix is a fixed paragraph appended to the unchanged base prompt. It asks the model to set aside the existing implementation strategy, describe the simplest implementation it would write from the observable contract and failing behavior, and then make the smallest change. The normal arm’s prompt is byte-identical to the control prompts of earlier P-series runs. 3 4

That block is the design restated for readability, not the frozen configuration, and the difference matters. The frozen P3 export records the two arms as {"name": "P3C", "models": ["qwen-local"], "samples": 12, ...} and {"name": "P3CF", ...} โ€” identical except for their names. The suffix text appears nowhere in the export. CodeAI’s ledger stores prompt_suffixes in the immutable experiment.created event, but the exporter never copied that field, so the one variable the experiment manipulated was dropped on the way out. 5 6

The rows still carry an indirect trace of it. On every one of the twelve tasks, every counterfactual call consumed exactly 44 more input tokens than every normal call on the same task. A fixed suffix produces precisely that pattern. It is consistent with the preregistered design; it is not the suffix text itself. 5

The exporter gap is repaired in the current CodeAI working tree: exported arms now include prompt_suffixes and stance_labels, exported call rows carry each call’s prompt_variant, and a regression test builds a normal-versus-counterfactual pair and asserts the export can tell them apart. The frozen P2 and P3 exports are unchanged and still lack the field. 6 7

The decision, then the interpretation

The rule’s inputs, recomputed from the frozen rows: 8

Normal ร— 12 Counterfactual ร— 12
Task coverage 9/12 (0.75) 9/12 (0.75)
Candidate passes 44/144 (0.306) 44/144 (0.306)
Input tokens 20,244 26,580
Output tokens 8,081 14,723
Total tokens 28,325 41,303 (+45.8%)
Median call latency 3.1 s 3.7 s

Solve sets were identical with zero rescues either way across 288 calls and 286 checks. 8 5

The primary hypothesis was higher verified coverage for counterfactual. Coverage held at 9/12, so the primary hypothesis fails. The preregistered interpretation rules then apply as written. Rule 2 โ€” counterfactual approximately equal to normal โ€” reads the P2 difference as sampling noise and keeps normal as the default. Rule 6 fires alongside it: retry-once-accepted and size-format remain unsolved in every arm, so framing variants stop being aimed at them. 1 9

One sentence in the P3 report goes further than the rules: that the token premium alone would reject counterfactual as a default. The preregistered rules contain no token threshold, so that is the report’s judgment rather than a frozen clause. It leaves the decision unchanged here, because the coverage tie already decides, but a reader applying the same method should put the resource multiple into the rule before the draws, as Dodge’s argument implies. 9

The premium itself has a mechanism the rows can show. Of the 12,978 extra tokens, 6,336 are the suffix (44 tokens ร— 144 calls) and 6,642 are longer outputs: counterfactual responses were longer on all twelve tasks. The prompt cost more than its own length; it changed how much the model wrote. 5

Keeping the default is not proving normal superior. The challenger failed the condition for replacing the incumbent in this test โ€” one local model, twelve tasks, one wording โ€” and nothing more.

The tie that moved

Equal aggregates do not mean identical rows. Per task, recomputed from the frozen export and matching the report cell for cell: 5

Task Normal passes / 12 Counterfactual passes / 12
unsound-cache 3 9
dt-roundtrip 10 6
env-timing 10 8
lsp-square 6 8
half-up-rounding 6 5
single-append 1 2
splitlines-cr 3 2
stable-priority 4 3
registry-pollution 1 1
retry-once-accepted 0 0
size-format 0 0
strip-query 0 0

Both columns sum to 44 over the same nine covered tasks, yet the passes sit in different places: up six on one task, down four on another, smaller shifts both ways. On this run, counterfactual wording changed where success appeared without improving the metric required for promotion. That is the accurate summary, and it rules out the lazy one: “the wording did nothing” is not what the rows show.

It is fair to ask the next question, though: is that movement more than draw-to-draw noise? Nothing was preregistered to answer it, so what follows is a post-hoc check, run over the frozen rows while editing this chapter, labeled as exactly that. If wording had no effect, each task’s passes would be exchangeable between the two arms. Shuffle each task’s 24 draws between the arms, many times, and see how often the total movement is as large as the observed 18:

import random

rows = [(3, 9), (10, 6), (10, 8), (6, 8), (6, 5), (1, 2),
        (3, 2), (4, 3), (1, 1), (0, 0), (0, 0), (0, 0)]   # (normal, counterfactual)
observed = sum(abs(a - b) for a, b in rows)                 # 18

def shuffled_movement():
    total = 0
    for a, b in rows:
        draws = [1] * (a + b) + [0] * (24 - a - b)
        random.shuffle(draws)
        x = sum(draws[:12])
        total += abs(x - (a + b - x))
    return total

trials = 200_000
p = sum(shuffled_movement() >= observed for _ in range(trials)) / trials
# p โ‰ˆ 0.31

About 31% of shuffles move success at least as much as the real arms did, and while the single largest shift โ€” unsound-cache at 3 against 9 โ€” gives an uncorrected Fisher exact p of about 0.04, it is one of eight tasks whose counts differ, so it survives no reasonable correction for looking at all eight. The P3 report reached the same reading in words: task-by-prompt interaction noise, not a stable framing effect. 5 9

So the movement happened, and twelve draws per task cannot distinguish it from chance, which is why it stays unpromotable. The chapter claims observed redistribution only โ€” not specialization, not a task-wording affinity, not a routing rule. The reading rule for ties is symmetric: unfold the aggregate before concluding that nothing happened, and test the unfolded rows before concluding that something did.

The two missing checks

Two hundred eighty-eight calls produced 286 checks. The frozen rows identify both gaps: one EXECUTION_ERROR candidate per arm with no check ID โ€” a normal draw on stable-priority and a counterfactual draw on strip-query. Candidates that fail in execution never reach the hidden tests, the same accounting the P1.1 report states for its compile-gate errors. Coverage arithmetic is unaffected, because an execution error is a failed draw either way. 5

A second quirk stays logged rather than normalized: every P3 call row, like P2’s, labels its prompt version seeded-code-v1 on a semantic-repair-v1 run. It affects no arm identity, task mapping, or metric, and it is one more reason the suffix trace above matters โ€” the version field could not have told the arms apart either. 5

Checking it without trusting it

The analysis verifier asserts the P3 identities โ€” 9/12 against 9/12, 44/144 against 44/144, identical solve sets, zero rescues โ€” and catches its seeded corruption; the export verifier checks all four frozen bundle hashes. Those cover the aggregates the rule consumes. 8

Everything else in this chapter is recomputation over the same frozen rows, done for this chapter and labeled as such: the per-task table, the input/output token split, the constant 44-token input delta, the execution-error identification, and the permutation check. None of it is a new bundle, and none of it required a model call. 5

The report’s limits are kept whole: one local model, twelve tasks, one counterfactual wording, within-corpus replication, no out-of-sample validity. With no effect to carry outward, the planned out-of-sample follow-up is moot by the rules’ own logic. Declining to fund another replication is a resource decision the evidence supports. 9

The same instincts show up outside laboratories. In one self-reported public thread, an author running AI-rewritten trading rulebooks against a static control found most AI accounts trailing it and called their best account luck; a reply named the multiple-comparisons problem (thread). It evidences nothing here. It is a reminder that a control and a suspicion of the winner are cheap, and usually available. 10

What this is not

  • Not proof the P2 signal was noise. The rule’s output is that the challenger did not earn promotion, and Rule 2 reads that outcome as noise; the mechanism behind P2’s difference remains unidentified.
  • Not a vindication for normal. One incumbent survived one test, nothing broader.
  • Not evidence the wording did nothing. Success moved between tasks without improving the promotion metric, though whether the wording caused that movement is not established.
  • Not out-of-sample. Nothing here travels to other tasks, models, or wordings.

Where it is still weak

  1. Twelve tasks, one model, one wording. Significance and generality are unavailable. 9
  2. Movement without a test that could see it. Twelve draws per task cannot separate redistribution from chance; the post-hoc check was not preregistered. 5
  3. The manipulated variable is not in the frozen export. The suffix is attested by the preregistration, the report, and a constant input-token delta, not by a recorded field. 5
  4. Checks trail calls by two. Both gaps are identified execution errors. 5
  5. The token rejection is not a frozen clause. The rules had no resource threshold. 1

Do this now

Thirty minutes. Run one replication whose answer could embarrass you.

  1. Take an attractive subgroup finding of yours. Write the matched question it would need to survive: same exposures, same metric, one variable changed.
  2. Write the decision rule before collecting anything: what margin promotes, what retains, what drops โ€” and the resource multiple at which a tie still loses.
  3. Check that your export records the variable you are changing. If two arms export identically except for their names, fix that first.
  4. Run it. Score the rule in one sentence before any per-task inspection.
  5. Then unfold the aggregate task by task, and run a shuffle test like the one above before you believe any single row.

If you are building with an assistant:

Give every promising signal a matched replication before promotion: same
draw counts, same tasks, one variable changed, rule frozen first, including
any resource threshold. Verify the exported configuration records the
changed variable. Score the rule before interpreting; a tie at higher cost
does not promote. Then unfold aggregates task by task, report movement in
both directions, and test it against a shuffle null before describing it as
an effect. Account for every missing check by row and rename no historical
label.

Failure modes

  • Calling the tie a near-win. Same coverage at +46% tokens fails the promotion condition.
  • Claiming nothing happened. Success moved; totals hide it.
  • Claiming something promotable happened. Movement that a shuffle reproduces a third of the time is not a routing rule.
  • Adding the threshold afterward. A resource clause invented after the draws is judgment, not preregistration.
  • Trusting the export’s arm names. If the changed variable is not in the record, the record cannot show the experiment was the one you meant.

What this chapter established

  • The matched replication tied on the rule’s metrics: 9/12 coverage and 44/144 passes each, identical solve sets, zero rescues, at +45.8% tokens โ€” half suffix input, half longer outputs. 8
  • Under the frozen rules, normal remains the default, and framing attacks on the two universally unsolved tasks stop. 1
  • Counterfactual wording changed where success appeared without improving the promotion metric; a post-hoc shuffle check (p โ‰ˆ 0.31) shows twelve draws per task cannot distinguish that movement from chance. 5
  • The frozen export omits the manipulated prompt suffix; a constant 44-token input delta is the rows’ only trace of it, and the exporter is repaired in current source with a regression test. 5 6

Next

Generation policy now has its discipline: signals earn matched tests, tests earn decisions only through frozen rules, and rules may keep the incumbent. But every experiment in this part spent model calls freely in order to ask its question. The remaining question is whether spending another call should itself be governed โ€” not which model answers, but whether any model needs to be asked at all.

Continue with What Should Happen Next?.

References

  • Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven, and David Thomas Mellor. The Preregistration Revolution. Proceedings of the National Academy of Sciences 115(11):2600โ€“2606, 2018. PNAS.
  • Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. Show Your Work: Improved Reporting of Experimental Results. EMNLP-IJCNLP, 2019. ACL Anthology.

Implementation sources: P3 ran on the P-series CodeAI lineage (src/codeai/stances.py: stance_suffix, stance_prompt; src/codeai/experiments.py: ArmDef, run_arm; src/codeai/analysis.py: arm_metrics, unique_rescues, build_report, export_experiment). The exporter repair (arm prompt_suffixes and stance_labels; per-call prompt_variant) and its regression test tests/test_stances.py::test_export_carries_the_prompt_variable_of_each_arm are uncommitted in the CodeAI working tree on top of a1b562a; the full suite passed (451). The ArmDef block restates the preregistered design in the current API; it is not the frozen configuration. The permutation check was executed over frozen rows during editing and is not preregistered. Evidence: experiments/applied-ai/evidence/p-series/ (frozen exports, hashes verified), experiments/applied-ai/evidence/p-series-analysis/ (analysis.json, TABLES.md, analyze_pseries.py, verify_analysis.py), C:/Projects/codeai/experiments/P3-prereg.md and P3-results.md (frozen preregistration and report). Nothing was rerun and nothing frozen was modified.


  1. Report: C:/Projects/codeai/experiments/P3-prereg.md↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Report: C:/Projects/codeai/experiments/P2-results.md↩︎

  3. Source inspection: src/codeai/stances.py (stance_suffix). ↩︎

  4. Source inspection: src/codeai/experiments.py (ArmDef). ↩︎

  5. Measured run: experiments/applied-ai/evidence/p-series/p3↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  6. Source inspection: src/codeai/analysis.py (export_experiment). ↩︎ ↩︎ ↩︎

  7. Source inspection: tests/test_stances.py (test_export_carries_the_prompt_variable_of_each_arm). ↩︎

  8. Measured run: experiments/applied-ai/evidence/p-series-analysis↩︎ ↩︎ ↩︎ ↩︎

  9. Report: C:/Projects/codeai/experiments/P3-results.md↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  10. Report: docs/applied-ai/plans/capstone-selective-intelligence.md↩︎