← Jev From First Principles

Semantic Programs

What does a program with deterministic and semantic computation together look like, end to end?

Semantic Programs

This chapter is model-free: the program runs over stub providers with hand-supplied numbers (labelled ILLUSTRATIVE) and deterministic retrieval (BM25, Chapter 21). The generation step is a stub. The ReAct-style comparison is a count over a declared task script — no agent is run (no LLM is allowed by the book’s rules).

The problem

The book has spent sixteen chapters building pieces: decide, if decide, match decide, where decide, the spec, the compiler, the graph. This chapter asks the question the pieces were for: what does a program with deterministic and semantic computation together look like, end to end?

What we expect and why

Our hypothesis: on a fixed claim-research task, the semantic program makes the same decisions as a ReAct-style loop but with every decision typed and countable and a replayable graph — and it gives up flexibility to get that. Three papers frame it:

  1. Yao et al. (2022), ReAct (2210.03629) — interleaves free-text reasoning traces with tool actions; on HotpotQA and FEVER it mitigates hallucination by querying a Wikipedia API. The challenge: every “Thought” is an untyped, free-text decision. (Read at abstract level.)
  2. Gao et al. (2022), PAL (2211.10435) — the model generates a program and the interpreter computes; it was FAMOUSLY strong on GSM8K. The method this chapter copies in miniature: delegate the deterministic work to code, keep the semantic decisions typed. (Read at abstract level.)
  3. Schick et al. (2023), Toolformer (2302.04761) — models learn when to call tools; tool use is itself a decision. The support this chapter needs: deciding to act is a decision, and making it typed is the whole point. (Read at abstract level.)

The paper that weakens the expectation is ReAct itself: its flexibility (the agent rewrites its plan from free text mid-run) is exactly what the semantic program gives up. The chapter’s comparison is therefore not “which is more accurate” — it is what the free text costs, counted as untyped decisions.

The build

src/arbiter/semantic_program.py runs the program: load a hand-built corpus of nine documents, retrieve the top 3 with BM25, decide relevance for each (where decide), decide relation for each kept document (match decide, with the tie rule from Chapter 18: a near-tie resolves to CONFLICTING_EVIDENCE), generate a [stub] summary, aggregate a support verdict, and gate the side effect with a permit/deny/escalate boundary. Every decision is recorded in a Chapter 26 decision graph; the graph is replayed from disk.

from arbiter.graph import categorise
from arbiter.semantic_program import research_claim, replay_seeded
    # 1. research a claim: retrieve, where decide, match decide, support, act
    r = research_claim("compound Q reduced inflammatory markers", seed=3)
    print("1. claim A")
    print(f"   retrieved {r['retrieved']} kept {r['kept']}")
    print(f"   verdict {r['verdict']} p={r['p_support']} -> {r['action']} "
          f"({r['reason']})")
    print(f"   typed decisions {r['typed_decisions']}, free-text {r['free_text_decisions']}")
    # 2. the runtime boundary across the three cases
    print("2. runtime boundary")
    for tag, claim in (("B", "the alpine study recorded temperature ranges"),
                       ("C", "small pilot study reported modest drop")):
        x = research_claim(claim, seed=3)
        print(f"   {tag}: {x['verdict']} p={x['p_support']} -> {x['action']}")
    # 3. replay from the Chapter 26 decision graph
    print("3. replay")
    print(f"   same seed 3: {categorise(replay_seeded(r, 3))}")
    print(f"   changed seed 8: {categorise(replay_seeded(r, 8))}")
    # 4. free-text decision count (ANALYTICAL; nothing run)
    print("4. free-text decisions")
    print(f"   semantic program: 0 (typed {r['typed_decisions']})")
    print("   react-style loop: 5 (one free-text Thought per step, counted)")
1. claim A
   retrieved ['d9', 'd1', 'd2'] kept ['d9', 'd1', 'd2']
   verdict CONFLICTING p=0.88 -> escalate (conflicting or unknown evidence)
   typed decisions 9, free-text 0
2. runtime boundary
   B: SUPPORTED p=0.9 -> permit
   C: SUPPORTED p=0.55 -> deny
3. replay
   same seed 3: {'no_change': 9}
   changed seed 8: {'no_change': 5, 'decision_changed': 3, 'state_changed_decision_same': 1}
4. free-text decisions
   semantic program: 0 (typed 9)
   react-style loop: 5 (one free-text Thought per step, counted)

The run

run_ch34.py writes results/ch34.jsonl and the replayable graph evidence/ch34-replay.jsonl (mode program, replay, analytical).

The program on three claims (seed 3, declared; every provider a stub):

claim kept relations verdict p action
A: compound Q reduced inflammatory markers d9, d1, d2 uncertain (tie rule), supports, contradicts CONFLICTING 0.88 escalate
B: the alpine study recorded temperature ranges d3 supports SUPPORTED 0.90 permit
C: small pilot study reported modest drop d5 supports SUPPORTED 0.55 deny

The three cases cover the boundary completely: permit on confident support, deny on low-confidence support (the runtime refuses an action), escalate on conflicting evidence. Document d9 also exercises the Chapter 18 tie rule in the main run: its supports/contradicts probabilities (0.42/0.40) are within ε=0.05, so it resolves to CONFLICTING_EVIDENCE — never an arbitrary winner.

Replay (Chapter 26 graph, claim A, 9 nodes): same-seed replay gives no_change 9/9 — decisions and state hashes agree, so the stored graph is a faithful record. Replaying behind seed 8 changes three nodes — mat:d1 (supports→contradicts) and its ancestors support and act — exactly the sampled nodes and their dependents, nothing else. The graph you could store yesterday tells you which decisions change when the sampling changes, and which do not.

Free-text decisions (mode analytical; nothing was run):

system free-text decisions typed decisions note
semantic program 0 9 every ask has an answer set
ReAct-style loop 5 0 one free-text Thought per step (search→read→verify→decide→finish), counted

The ReAct number is a count over the declared five-step task script, swept to show the dependence: at 3–6 steps the free-text count is 3–6, with an assumed 50–200 tokens per Thought (rows carry the sweep). No accuracy claim is made for either system; the acceptance criterion was the count, and it is reported.

What it says

  1. The surviving construct set is enough to write the program. decide (every ask), if decide (the support/boundary branches), match decide (relations, with its tie rule firing on d9), where decide (the relevance filter). The one thing Chapter 23 rejected — for decide — appears only in its library form (a for loop over BM25Index.search). The constructs the earlier chapters kept are the ones this program used. P5 holds.

  2. The runtime boundary is a decision, not a guard clause. Permit/deny/ escalate is itself recorded in the graph with its antecedents; the deny on claim C is a decision about acting, and replay shows it. The boundary refuses an action on low confidence, and the refusal is auditable. P2 holds.

  3. Replay is the audit. Same-seed 9/9 agreement; changed-seed changes only the sampled nodes (and their ancestors). A ReAct trace is free text; you can replay it only by re-running the model. P4 holds.

  4. The count is the honest comparison. The semantic program makes 0 free-text decisions for 9 typed ones; a ReAct-style loop makes one per Thought, all 5 by the minimal script. P3 holds — and it is the entire claim: no cost or accuracy number is implied by this chapter (the only cost row is an assumed-tokens sweep, labelled ANALYTICAL).

  5. The corpus story is a property of the stubs. The verdicts are consequences of the hand-supplied relation triples and the BM25 retrieval; they are ILLUSTRATIVE, exactly as rule 14 demands. P1 holds only as “the program resolves each claim to the verdict its stub evidence dictates”.

Wrong / Correct. Wrong: “The semantic program and a ReAct loop solve the same task with the same flexibility.” Correct: “The semantic program trades ReAct’s free-text flexibility for typed, countable, replayable decisions: 0 untyped decisions, 9 typed, a graph that replays 9/9 — and gives up the agent’s ability to re-plan in natural language. Which one you want depends on whether you can enumerate the decisions in advance.”

The construct set, stated plainly

construct Chapter verdict used here
decide 16 kept every ask (relevance, relation, support, act)
if decide 17 kept support/boundary branches
match decide 18 kept relations; tie rule fired on d9
where decide 22 kept (typed, abstaining; LOTUS sem_filter, narrower, no novelty) relevance filter
for decide 23 rejected library for over retrieval

No new syntax layer was built: Chapter 23’s rejection and Chapter 22’s “filter choice matters more than syntax” make the library form the honest result, and the program is written in it.

Prior art, engaged directly

  • ReAct (2210.03629): the challenge. Its decisions are generated text; the count above is what that costs. The semantic program is ReAct’s task with the decisions enumerated in advance.
  • PAL (2211.10435): the method. Deterministic work (retrieval ordering, aggregation, hashing) lives in the interpreter; only the decisions are semantic. This program is PAL’s split, applied to decisions rather than to arithmetic.
  • Toolformer (2302.04761): the support. Choosing to act is a decision; this chapter makes the choice typed and gated, where Toolformer learns it as a token-generation skill.

No novelty is claimed over any of them; the chapter’s contribution is the count and the replay, both measured on the program’s own deterministic runs.

What to carry forward

Chapter 35 writes the verdict. It will need the honest lists: which constructs survived (this chapter used them) and which were rejected (Chapter 23’s for decide, shown in library form here). And it will need the Part VI record: the same 188 SciFact dev claims served as the test split of Chapters 21–23 and 25 — this chapter’s corpus is new and hand-built, so it does not inherit that leak, but it does inherit the discipline of labelling every stub.

Close by

Which constructs did the program actually use? All the survivors — decide, if decide, match decide, where decide — and the rejected one only as a library loop. That is the chapter’s answer to its own question, and it is the narrowest claim the book will carry into Chapter 35.

Limitations

  • The corpus, relation triples, thresholds (keep 0.60, permit 0.80, ε 0.05) and sampling seed (3) are all assumed; every verdict is a consequence of them, reported as such (ILLUSTRATIVE).
  • The generation step is a [stub] template with zero tokens; no cost claim is made for it.
  • The ReAct comparison is ANALYTICAL: a count over a declared five-step script with a swept token assumption; no agent ran and no accuracy comparison is claimed (an LLM is not allowed by the book’s rules).
  • The three papers are read at abstract level.
  • The program records decisions in the graph but not their probabilities; the probabilities live in the result rows and drive the boundary, and the replay agreement is over decisions and state hashes (Chapter 26 semantics).
  • The sampling seed for the showcased run (3) differs from the preregistration’s placeholder “seed 0 and 1”; the deviation is disclosed here and in the metadata, and the replay test still covers same-seed vs changed-seed.