← Jev From First Principles

The First Token

What tokenisation does to option scoring: which readouts exist at 77 options, which are impossible, and what that does to the claim that a decision model is just a way of asking a model.

Chapter 4 fixed a world with 77 labels and put a number on it. Chapter 5 loosened one joint: the label set could be described at run time instead of trained in. Both chapters compared providers on accuracy. This one asks a question that comes before any provider runs.

Here is the shortest version of it. A decision model that returns a probability over a set of options must, somehow, turn “which of these?” into a number. The most common way is to write the options into a prompt and read the model’s distribution over whatever comes next. If the options are printed as letters, that is a probability over 26 possible answers. If they are printed as words, it is a probability over strings of different lengths. If they are printed as numbers, it is a probability over digits — and digits are not the same size in every tokenizer. The third-party account of how Jev works that this chapter set out to rebuild (research/primary-sources.md, Victor Dibia’s write-up — not vendor documentation; the vendor has published no internals) describes exactly this: a prompt that ends where the answer goes, then each option scored by its log-probability, then a softmax across options. That is the mechanism. It is also, as we are about to see, at least four different mechanisms depending on how the options are printed.

The claim under test is the book’s own: H1, that decision models reduce to known techniques plus packaging. If “option scoring” turns out to be one thing, H1 wins its easiest clause. If it turns out to be a family whose members disagree about whether they exist at all at 77 options, then the clause was never a claim about one thing, and the useful question becomes which member you picked and why.

What we expected, and why

Robinson & Wingate (arXiv:2210.12353) are the source of the technique this chapter builds: present the options as symbol-enumerated lines and read the probability of the symbol at the answer position. Their §3 lists four problems with the older alternative — scoring each option’s text separately — and the third is the one that matters here: with cloze-style scoring you must choose a normalisation, and that choice “often incur[s] a computational cost or depend[s] on choice of tokenization scheme”. Their Table 3 measures it: randomly re-casing or re-spacing the option text costs the cloze approach 12.4% and 10.3% accuracy, and costs the letter approach 1.3% and 0.5%. Surface form is a big deal for one readout and nearly irrelevant for the other. That is a strong hint that the printing of the options is not a formatting detail.

Zheng et al. (arXiv:2309.03882) qualify the technique in a way that a later chapter will have to fight: the bias that makes option order matter is, in §2.4, mostly token bias rather than position bias — the model a priori assigns more mass to some ID tokens — and removing the IDs reduces it while degrading accuracy. Their §2.4 also reports that swapping the symbol alphabet (a/b/c/d, 1/2/3/4, (A)/(B)/(C)/(D)) does not help. If the choice of symbol does not matter for bias, then our chapter’s question is somewhere else. Zhao et al. (arXiv:2102.09690) supply the correction the third-party account assumes: divide out a prior estimated from a content-free input.

Then, in the last twelve months, the prior art moved. LLM-as-Jev (arXiv:2610.02076) presents bracketed numeric identifiers [1] … [K] and argues in §3.2 that they are prefix-free because every suffix ends in ], so the option count is unbounded where letters stop at 26. Its Appendix C observes that Qwen tokenises digits individually, so “options 1 and 10 share their first token” — and its Appendix H reports a benchmark where fine-tuning moved answers from option 10 to option 1, with the honest admission “We have not identified the cause”. Four searches of arXiv (recorded with dates in evidence/notes-ch07.md) found no paper that measures first-token collision rates over a real label set. So the measurement below is not a reproduction; it is the missing measurement.

So we predicted, and committed before running (metadata/07-chapter.yaml, commit a19ff6b), twelve token-level predictions. The honest summary of what we expected: letters would be collision-free and would stop existing above 26 options; bare digits 1–77 would be both non-prefix-free and collision-heavy; bracketed identifiers would be prefix-free but otherwise unchanged; and the 77 label names would collide heavily enough to block a first-token readout of the strings. Six held. We report the six refutations with the same prominence.

The build

src/arbiter/providers/option_scoring.py, written from scratch on the Chapter 2 contract. Nothing was forked. It has two readouts, named for what they are rather than for what they are supposed to be:

  • letter readout — the probability of a single-token symbol at the answer position, restricted to the option symbols and renormalised. Robinson & Wingate’s MCP; AnyJev’s raw. One token per option, so one forward pass.
  • full-option-string log-probability — the sum of the option’s token log-probabilities, softmaxed across options. Robinson & Wingate’s CP; LLM2Jev’s Eq. (1).

Two decisions were declared before any run and are visible in the code. The multi-token rule is the sum, not the mean, because a sum is the joint likelihood of strings of unequal length and a mean is not a probability of anything; the cost is that the score structurally prefers short options. (The same confound appears independently in 2607.27421’s calibration section, which warns that the sequence-level log-probability it uses is “sensitive to output label length”.) And the letter readout is capped at 26 options, raising at fit rather than at the three-thousandth item — a ceiling we took from AnyJev’s own code, MAX_OPTIONS = 26, because 26 is where the letters run out and a provider that truncates a 77-way question is worse than one that refuses it.

The tokenisation finding

Run now, on this machine, CPU only, tokenizer only. 6001 rows in results/ch07.jsonl, all tagged "mode": "observed-cpu". Three pinned tokenizers: Qwen3-1.7B (instruct), Qwen3-1.7B-Base, and Llama-3.2-1B-Instruct as a second family. Twelve option sets (the 77 label names, four Chapter 5 wordings of all 77, the 17 held-out labels under four wordings, the two safety labels under four wordings) and six ways of printing them.

The script below produces that table. It reads the committed results/ch07.jsonl, so it runs in a fraction of a second and needs no model: the tokenizers were the only thing consulted when the rows were written.

"""The chapter's code block: the tokenisation ledger, computed from
results/ch07.jsonl. No model, no GPU, no network, under a second.

    python examples/ch07-the-first-token/demo_ch07.py

Everything printed is read out of the results file or recomputed from it. If a
number here is not in that file, this script is wrong and the chapter is wrong
with it; that is the point of generating the block rather than typing it.
"""

from __future__ import annotations

import json
import sys
from collections import defaultdict
from pathlib import Path

ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT))
sys.path.insert(0, str(ROOT / "src"))

RESULTS = ROOT / "results" / "ch07.jsonl"
PRIMARY = "qwen3-1.7b-instruct"
SECOND = "llama-3.2-1b-instruct"

rows = [json.loads(l) for l in RESULTS.open(encoding="utf-8")]


def cell(tokenizer: str, option_set: str, presentation: str,
         metric: str) -> float | None:
    for r in rows:
        if (r.get("provider") == f"tokenizer/{tokenizer}"
                and r.get("option_set") == option_set
                and r.get("presentation") == presentation
                and r.get("metric") == metric):
            return r["value"]
    return None


def infeasible(tokenizer: str, option_set: str, presentation: str) -> bool:
    for r in rows:
        if (r.get("provider") == f"tokenizer/{tokenizer}"
                and r.get("option_set") == option_set
                and r.get("presentation") == presentation
                and r.get("metric") == "presentation_feasibility"):
            return r["value"] == 0.0
    return False


def depth_groups(tokenizer: str, option_set: str, presentation: str,
                 d: int) -> float | None:
    for r in rows:
        if (r.get("provider") == f"tokenizer/{tokenizer}"
                and r.get("option_set") == option_set
                and r.get("presentation") == presentation
                and r.get("metric") == "exploratory_prefix_groups_at_depth"
                and r.get("read_depth") == d):
            return r["value"]
    return None


def verdict(pid: str) -> tuple[str, str]:
    for r in rows:
        if r.get("metric") == f"prediction_{pid}":
            return r["verdict"], r["value"]
    return "MISSING", -1.0


print("JEV-07-02  what tokenisation does to option scoring")
print(f"tokenizer: Qwen/Qwen3-1.7B @ 70d244cc (primary), "
      f"meta-llama/Llama-3.2-1B-Instruct @ 92131767 (second family)")
print(f"option sets: 77 BANKING77 label names, 4 wordings x 77, "
      f"17 held-out x 4, 2 safety x 4")
print(f"rows in results/ch07.jsonl: {len(rows)}\n")

print("77 options, one presentation per line")
print("  presentation    readout           distinct  collide  pairs  "
      "pfx_tok  depth_to_separate")
for pres, name in (("letter", "letter readout"),
                   ("number_bare", "bare digits"),
                   ("number_bracketed", "bracketed [k]"),
                   ("full_string", "full option strings")):
    if infeasible(PRIMARY, "banking77_names_w0", pres):
        print(f"  {pres:15s} {name:17s} IMPOSSIBLE by construction "
              f"(26 letters < 77 options)")
        continue
    d1 = cell(PRIMARY, "banking77_names_w0", pres, "distinct_first_tokens")
    col = cell(PRIMARY, "banking77_names_w0", pres,
               "options_first_token_collision")
    pairs = cell(PRIMARY, "banking77_names_w0", pres,
                 "first_token_collision_pairs")
    pfx = cell(PRIMARY, "banking77_names_w0", pres, "strict_prefix_pairs_token")
    depth = cell(PRIMARY, "banking77_names_w0", pres,
                 "exploratory_max_depth_to_separate_all")
    depth_s = "1" if depth == 1 else f"{depth:.0f}"
    print(f"  {pres:15s} {name:17s} {d1:8.0f} {col:8.0f} {pairs:6.0f} "
          f"{pfx:8.0f}  {depth_s}")

print("\n17 held-out options and 2 safety options, same tokenizers")
print("  set            pres              distinct  collide  pairs  pfx_tok")
for oset in ("unseen17_w0", "safety_w0"):
    for pres in ("letter", "number_bare", "number_bracketed", "full_string"):
        if infeasible(PRIMARY, oset, pres):
            print(f"  {oset:14s} {pres:17s} IMPOSSIBLE")
            continue
        print(f"  {oset:14s} {pres:17s} "
              f"{cell(PRIMARY, oset, pres, 'distinct_first_tokens'):8.0f} "
              f"{cell(PRIMARY, oset, pres, 'options_first_token_collision'):8.0f} "
              f"{cell(PRIMARY, oset, pres, 'first_token_collision_pairs'):6.0f} "
              f"{cell(PRIMARY, oset, pres, 'strict_prefix_pairs_token'):8.0f}")

print("\nsurface form moves the first token (77 label names, Qwen3)")
print(f"  one leading space changes the first token of "
      f"{cell(PRIMARY, 'banking77_names_w0', 'full_string', 'leading_space_changes_first_token'):.0f} of 77")
print(f"  capitalising the first letter changes it for "
      f"{cell(PRIMARY, 'banking77_names_w0', 'full_string', 'capitalisation_changes_first_token'):.0f} of 77")
lens = [cell(PRIMARY, "banking77_names_w0", "full_string", k) for k in
        ("token_len_min", "token_len_median", "token_len_max")]
print(f"  option token length min/median/max: {lens[0]:.0f}/"
      f"{lens[1]:.0f}/{lens[2]:.0f}")

print("\nwhich presentation collides, per wording (77 options)")
print("  wording   distinct  collide  largest class")
for w in ("w0", "w1", "w2", "w3"):
    oset = "banking77_names_w0" if w == "w0" else f"banking77_{w}"
    print(f"  {w:8s} "
          f"{cell(PRIMARY, oset, 'full_string', 'distinct_first_tokens'):8.0f} "
          f"{cell(PRIMARY, oset, 'full_string', 'options_first_token_collision'):8.0f} "
          f"{cell(PRIMARY, oset, 'full_string', 'largest_first_token_class'):14.0f}")

print("\ntwo tokenizer families, bare digits (77 options)")
for tok, label in ((PRIMARY, "Qwen3  "), (SECOND, "Llama3.2")):
    print(f"  {label} distinct first tokens: "
          f"{cell(tok, 'banking77_names_w0', 'number_bare', 'distinct_first_tokens'):.0f}"
          f"   strict prefix pairs: "
          f"{cell(tok, 'banking77_names_w0', 'number_bare', 'strict_prefix_pairs_token'):.0f}")

print("\nthe largest colliding first-token classes (Qwen3, w0 label names)")
for r in rows:
    if (r.get("option_set") == "banking77_names_w0"
            and r.get("presentation") == "full_string"
            and r.get("metric") == "largest_first_token_class_members"):
        print(f"  {r['note']}")
        break

print("\npredictions T1-T12, as preregistered")
holds = 0
total = 0
for i in range(1, 13):
    pid = f"T{i}"
    v, val = verdict(pid)
    if v == "NOT_APPLICABLE":
        continue
    total += 1
    holds += 1 if val == 1.0 else 0
    print(f"  {pid:4s} {v}")
print(f"  {holds} of {total} applicable predictions hold")

print("\nJEV-07-03  model run: DEFERRED")
print("  no accuracy, no calibration, no flip rate, no latency was measured.")
print("  commands in examples/ch07-the-first-token/STAGE_PLAN.md;")
print("  every model number in the chapter is PENDING_RUN.")
JEV-07-02  what tokenisation does to option scoring
tokenizer: Qwen/Qwen3-1.7B @ 70d244cc (primary), meta-llama/Llama-3.2-1B-Instruct @ 92131767 (second family)
option sets: 77 BANKING77 label names, 4 wordings x 77, 17 held-out x 4, 2 safety x 4
rows in results/ch07.jsonl: 6001

77 options, one presentation per line
  presentation    readout           distinct  collide  pairs  pfx_tok  depth_to_separate
  letter          letter readout    IMPOSSIBLE by construction (26 letters < 77 options)
  number_bare     bare digits              9       75    366       68  2
  number_bracketed bracketed [k]            1       77   2926        0  3
  full_string     full option strings       43       48     94        0  5

17 held-out options and 2 safety options, same tokenizers
  set            pres              distinct  collide  pairs  pfx_tok
  unseen17_w0    letter                  17        0      0        0
  unseen17_w0    number_bare              9        9     36        8
  unseen17_w0    number_bracketed         1       17    136        0
  unseen17_w0    full_string             14        6      3        0
  safety_w0      letter                   2        0      0        0
  safety_w0      number_bare              2        0      0        0
  safety_w0      number_bracketed         1        2      1        0
  safety_w0      full_string              2        0      0        0

surface form moves the first token (77 label names, Qwen3)
  one leading space changes the first token of 77 of 77
  capitalising the first letter changes it for 76 of 77
  option token length min/median/max: 2/3/8

which presentation collides, per wording (77 options)
  wording   distinct  collide  largest class
  w0             43       48             10
  w1             29       59             12
  w2             54       36              5
  w3             47       43             11

two tokenizer families, bare digits (77 options)
  Qwen3   distinct first tokens: 9   strict prefix pairs: 68
  Llama3.2 distinct first tokens: 77   strict prefix pairs: 0

the largest colliding first-token classes (Qwen3, w0 label names)
  'card':[7, 11, 12, 17, 18, 21, 33, 43, 47, 57]; 'top':[20, 26, 27, 30, 34, 42, 67]; 'transfer':[5, 28, 41, 74]; 'pending':[8, 16, 59, 62]; 'exchange':[3, 19, 52]

predictions T1-T12, as preregistered
  T1   HOLDS
  T2   REFUTED
  T3   REFUTED
  T4   REFUTED
  T5   REFUTED
  T6   REFUTED
  T7   HOLDS
  T8   HOLDS
  T9   HOLDS
  T10  REFUTED
  T11  HOLDS
  T12  HOLDS
  6 of 12 applicable predictions hold

JEV-07-03  model run: DEFERRED
  no accuracy, no calibration, no flip rate, no latency was measured.
  commands in examples/ch07-the-first-token/STAGE_PLAN.md;
  every model number in the chapter is PENDING_RUN.

Read the first block row by row, because the three surviving presentations fail in three different ways.

Letters do not exist at 77 options. Not “perform badly” — do not exist. There are 26 and the task needs 77. AnyJev’s code says the same thing (MAX_OPTIONS = 26), which is why that number is a constant in ours rather than a guess. This is the one row in the table that is a hard impossibility, and it is a property of the alphabet, not of any model.

Bare digits are a trap, and the trap is bigger than we predicted. Qwen3 gives 77 identifiers only 9 distinct first tokens: 1, 10, 11, 12 … 19 all start with 1, and so on, so 75 of the 77 options collide and there are 366 colliding pairs with a largest class of 11. There are also 68 strict token-prefix relations. A readout that reads one token cannot separate option 1 from option 11, and a readout that reads the full string has to decide what to do about the prefix nesting. Both failures come from the same fact: 10 is two tokens on this tokenizer.

Bracketing fixes the trap by creating a worse one. [1] … [77] has zero prefix relations, exactly as LLM2Jev §3.2 argues — the bracket is the trick, and it works. But every bracketed identifier now begins with the single token [, so the 77 options share one first token and all 77 collide. Reading depth 1 is not merely uninformative here, it is identically uninformative. Separation takes depth 3 on Qwen3 and depth 2 on Llama-3.2. The prefix-free design did not remove the collision; it moved it from a scattered pattern to a uniform one and charged one extra token of read depth for it. Our prediction T3 said the collision count would be unchanged. It went from 75 to 77. That refutation is the finding.

Full option strings collide less than we predicted, and are still blocked. We predicted at least 55 of 77 label names would share a first token; 48 do, with a largest class of 10 (the card … family: arrival, acceptance, linking, not working, payment fee, delivery estimate and so on). 48 of 77 is a majority, so a single-token readout of the strings is still unusable for this label set, but our threshold was too high and the direction was right. Wording moves the number a lot: the four Chapter 5 wordings give 43, 29, 54 and 47 distinct first tokens, so the same options presented four ways differ by 25 distinct tokens — more than the number of options lost to collisions in the best wording.

Surface form is not a detail. One leading space changes the first token of 77 of 77 label names. Capitalising the first letter changes it for 76 of 77. This is Robinson & Wingate’s Table 3 result (“Caps” and “Space” corruptions) arriving from the other direction: they measured that the cloze approach loses accuracy under those corruptions; we measured that the representation itself changes for almost every option, which is the mechanism behind their number. And the two tokenizer families agree exactly, on all four wordings — this is not a quirk of one vocabulary.

The finding that changes the most

The bare-digit trap does not exist on Llama-3.2. Llama tokenises 10 as a single token; Qwen3 splits it into 1, 0. So the same 77 options, printed the same way, under the same prompt, with the same readout, give:

  • Qwen3: 9 distinct first tokens, 75 colliding, 68 prefix pairs, separation at depth 2.
  • Llama-3.2: 77 distinct first tokens, 0 colliding, 0 prefix pairs, separation at depth 1.

LLM2Jev reports the Qwen behaviour as a fact about its tokenizer (Appendix C) and then finds a benchmark effect it cannot explain (Appendix H: answers moving from option 10 to option 1, “We have not identified the cause”). This chapter says what the cause is, at least structurally: on a digit-splitting tokenizer, options 1 and 10 are the same token until the second one. The effect is reproducible from the tokenizer files alone, with no model, no weights and about forty seconds of CPU — and it exists on one tokenizer family and not on the other. That is a stronger and much more actionable statement than “small models are biased”, and it is not in the literature we read.

What this does to “Jev is just option scoring”

Wrong: a decision model is an inference strategy — build the prompt, read the option probabilities, normalise — and what remains of Jev after you remove that is branding.

Correct: “read the option probabilities” is at least four different mechanisms. At 77 options one of them does not exist, one of them collapses 75 options onto 9 tokens, one of them collapses all 77 onto a single token, and the fourth needs up to 5 tokens of read depth — and which mechanism you get depends on which tokenizer you load, as the Qwen/Llama split above shows with everything else held constant. The interesting engineering is not “score the options”. It is choosing a presentation the tokenizer can read, refusing the ones it cannot, and knowing the read depth you are buying.

That is a claim about the shape of the family, and it is why this chapter ends up proposing a change to how the book states H1 (planning/pivot-proposals.md, 2026-10-07, status PROPOSED — the author decides). We did not edit planning/hypotheses.md.

The cost accounting (analytical, not measured)

One forward pass over the shared prompt, with the KV cache reused for every option, against the naive one-pass-per-option version: both are implemented in HFScorer (path="shared_prefix" and path="naive"), and the provider exposes the token counts through count_tokens. Those counts are ANALYTICAL and every row that carries them says so in its own accounting field. At 77 options the naive path re-reads the prompt 77 times; the shared path reads it once plus the option tokens. That is arithmetic. No wall-clock number has been measured for this provider and the chapter claims none.

PENDING_RUN: measured tokens and latency, both execution paths Command: python examples/ch07-the-first-token/run_ch07.py --stage S7 --metric tokens --device cuda then --stage S8 --metric latency --device cuda Fills: results/ch07.jsonl, rows with metric: tokens and metric: latency Note: the latency stage needs an idle machine, 200 items, warm, 3 repeats.

What the tests cover

python examples/ch07-the-first-token/selftest_ch07.py — 87 checks, no model, no GPU, about fifteen seconds.

  • A fake model with known logits. Every probability, normalisation, rotation, prior division and temperature is checked against arithmetic done by hand in the test file. The letter readout’s output is compared to a hand-computed softmax; the prior division to p/prior renormalised by hand; the batch prior to a running mean; the content-free prior to the mean of three probes. Nothing is trusted because it ran.
  • A trivially separable task (RUN.md standing lesson 2, which exists because Chapter 4 shipped a revision over an under-trained baseline). Six options with disjoint characters; the provider must get it right through the full DecisionRequest → ChoiceAnswer path, and it does.
  • The invariance that rotation has to preserve. Under an option-invariant fake, all cyclic rotations give zero flips and the same distribution; rotating and un-rotating the answer names one option across all rotations. Under a deliberately order-dependent fake, flips appear. Both directions are tested, because a rotation test that only ever passes proves nothing.
  • The mathematical claim behind log-space marginalisation. Under logit(i at position j) = c_i + b_j, the log-mean over all rotations returns softmax(c) exactly; the arithmetic mean does not. That is AnyJev’s argument and it is checked numerically rather than cited.
  • Two seeded bugs. (1) A tokenizer that merges the prompt’s last token with the continuation’s first — the trap LLM2Jev’s Appendix C warns about — must make the prefix invariant raise, and it does, both through the helper and through the provider. (2) The rotation permutation machinery is checked to be genuine: shift 0 is the identity, every shift is a true permutation, and every option visits every position exactly once — the properties whose absence would silently corrupt every answer while looking like an improvement.
  • A real code path. The last block builds a randomly initialised 2-layer Qwen3 in-process on the real Qwen3 tokenizer and runs the provider through it. This proves the transformers calls, the cache path and the prompt construction execute. Its numbers are meaningless by construction and are never written to results/.

Against the prediction

Clause by clause, six of twelve hold.

# Prediction Verdict
T1 bare digits, 77 options: 9 distinct first tokens, 75 colliding, 366 pairs, largest 11 HOLDS — exactly
T2 bare digits: 68 strict token-prefix pairs, 0 character-prefix pairs REFUTED — token half right (68), character half wrong: 1 is a character prefix of 10. Our own arithmetic error
T3 bracketed: 0 prefix pairs, collision count unchanged from bare REFUTED, substantively — 0 prefix pairs confirmed, but collisions go 75 → 77: the bracket makes all options share [
T4 letters: 26 distinct at K≤26, impossible at 77 REFUTED on the number, HOLDS on the substance — at 17 options there are 17 letters, not 26; we predicted the ceiling where we meant the count. Letters are collision-free (17/17 distinct, 0 pairs) and 77 is impossible
T5 label names: ≥55 of 77 share a first token, largest ≥15 REFUTED — 48 collide, largest 10. Still a majority, so the readout is still blocked; our threshold was too high
T6 label names: 0 character-prefix pairs, ≥1 token-prefix pair REFUTED — both zero. Sharing a first word is a collision, not a nesting: card arrival → [4951, 18647] and card not working → [4951, 537, 3238] diverge at the second token
T7 one leading space changes ≥70 of 77 first tokens HOLDS — 77 of 77
T8 capitalising changes ≥70 of 77 HOLDS — 76 of 77
T9 median 3–6 tokens, max ≥8, ≥95% multi-token HOLDS — median 3, max 8, 100% multi-token
T10 Llama-3.2 shows the same picture as Qwen3 REFUTED, and it is the chapter’s sharpest finding — Llama gives 77 distinct first tokens, 0 collisions, 0 prefix pairs
T11 safety labels: 2 distinct, no collisions or prefixes, all four wordings HOLDS
T12 unseen-17: 2–8 of 17 share a first token HOLDS — 6

Two of the six failures (T2, T6) were arithmetic and reasoning errors of ours, diagnosable in a few lines (examples/ch07-the-first-token/diagnose_refutations.py reproduces each). Two more (T4, T5) were the right direction with thresholds in the wrong place. Two (T3, T10) are findings about the world, and T10 is the one that changes the chapter’s conclusion.

The local run is deferred, and why

Local GPU work on this machine is paused by the author’s decision of 2026-10-07: five GPU runs had failed or been too slow, and a Chapter 6 run was still occupying the card. So Chapter 7’s model sweep did not happen. What that costs is specific, and it is the half of the chapter that would have said how often these structural limits are reached in practice.

What would change the book’s answer: if option scoring on frozen Qwen3 lands close to the Chapter 4 bar of 0.8779 on 77-way intent, H1’s “reduces to known techniques” clause is strongly supported and the packaging is most of it. If it lands far below, as the prior art suggests — 2607.27421 finds no model above 80% on Banking77 across a 41-model cohort, and LLM2Jev’s own frozen 4B reaches 69.0 — then “cheap” is the only thing option scoring has, and the interesting question becomes where in the accuracy-cost plane the intersection is. And if the readout collisions measured here turn out to predict accuracy across models, the readout’s availability is a first-order constraint on any system with a large label set, which is a stronger claim than H1 makes in either direction.

PENDING_RUN: all of JEV-07-03. Accuracy, macro-F1, ECE (15 equal-mass bins), Brier, AUROC, flip rate in option space, answer mass, tokens and latency, for Qwen3 at 0.6B / 1.7B / 1.7B-Base / 4B on 77-way intent, 17-way unseen and the two safety splits, raw and corrected, against the frozen Chapter 4 bar and the Chapter 5 embedding provider with paired bootstrap intervals. Command: stages S1–S9 in examples/ch07-the-first-token/STAGE_PLAN.md, one per model/task/split, each ≤15 minutes, each resumable, each writing a .done marker; merge with run_ch07.py --merge only when all nine markers exist. Fills: results/ch07.jsonl, then the results and prediction sections above. Estimated total: about 1 h 45 min of GPU time.

What we got wrong about our own predictions

We wrote twelve predictions expecting to be right about most of them and wrote them in a form precise enough to be wrong in a diagnosable way. Six held. Of the six that failed, two were arithmetic slips that took a minute each to find, two were threshold errors in the right direction, and two were discoveries about tokenisation that we could not have got right by reasoning alone — the bracketed scheme’s total depth-1 collision, and the fact that the digit trap exists on one tokenizer family and not another. That ratio is the argument for preregistering even the mechanical parts: the mechanical predictions were the ones that made the interesting failures legible, because a failed arithmetic claim is fast to diagnose and a failed intuition is not.

One more thing we would flag for the next chapter: the exploratory depth-of-read measurement (how many tokens before options separate) was not preregistered. It was added after T1 showed that bracketed identifiers share their first token, because a collision count with no statement of what to do instead is not actionable. It is labelled exploratory_ in every row, and it is what turned “there is a collision” into “you need depth 3, and here is the depth for each presentation”.

Limitations

  • No model ran. Every accuracy, calibration, flip-rate and latency number in this book for Chapter 7 is NOT_OBSERVED. The provider exists, is tested against a fake model with known logits, and has never been evaluated.
  • Three tokenizers, two families. Qwen3 (instruct and base) and Llama-3.2. Digit grouping is known to vary across families — Phi-4-style tokenizers group up to three digits, per LLM2Jev Appendix C — and we did not measure one. The family-dependence claim rests on two points, not a survey.
  • Two of the twelve predictions failed for our own arithmetic. We report the ratio honestly, but a preregistration with more slack would have predicted fewer failures and taught us less.
  • The label sets are one author’s. The four wordings come from Chapter 5’s frozen fixture, hand-written by the author without seeing dataset text. The collision counts are properties of these wordings; a different vocabulary would give different numbers, though the two tokenizer families agreeing exactly suggests the shape is robust to the tokenizer, not the author.
  • Only BANKING77 and the two safety labels. One 77-way set, one 17-way set, one binary set. 2610.07716 shows that the candidate menu itself moves accuracy 26 points on CLINC150, wider than a model-size step, so the menu is a first-order variable this chapter deliberately held fixed rather than varied.
  • The exploratory depth measurement is not preregistered. Labelled as such in every row; it is a lead, not a result.
  • No hosted Jev. src/arbiter/providers/hosted_jev.py does not exist in this repository and there is no response cache, so the optional comparison of Jev’s behaviour on the tokenisation-collision cases — does Jev separate the ten card … labels that a single token cannot? — is skipped, not deferred quietly. It is the most interesting question this chapter raises for a reader with an API key.

What Chapter 7 leaves behind

  • src/arbiter/providers/option_scoring.py — two readouts, declared multi-token rule, log-space rotation, batch and content-free priors, 26-option refusal, both execution paths, ANALYTICAL token accounting. Untested against a real model, and labelled as such in its own docstring.
  • examples/ch07-the-first-token/ — the tokenisation analysis (the chapter’s one real result), the 87-check selftest with two seeded bugs, the demo that prints this chapter’s code block, the refutation diagnosis, and the nine-stage plan for the deferred run.
  • results/ch07.jsonl — 6001 rows, "mode": "observed-cpu".
  • A dated entry in planning/pivot-proposals.md proposing that H1’s list of constituent techniques stop treating “first-token LM inference” as one thing.

The prompt ends with a question, and this chapter can only answer half of it. Is this a model, or a way of asking a model? The tokenisation answer is clear: at 77 options, whether you have a way of asking at all is settled by the alphabet and the tokenizer, not by the model. What would have to be learned for it to be a model? Nothing in this chapter tests that — the deferred run is the test. Chapter 8 takes the other half, which we can answer now: whatever the mechanism is, the program should not be able to ask for a decision the provider cannot represent, and the refusal in fit is the smallest honest version of that.