← Jev From First Principles

Semantic Match

Match on a decision with exhaustive arms, honest ties, and a measured answer to whether the wording moved the outcome.

Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.

if decide gave you two arms: value and uncertain. But decisions have more than two shapes. A choice among seventy-seven intents is not a boolean, and collapsing it to “answered or not” throws away the structure the caller actually branches on: which answer, how close the runner-up was, and whether the question was answerable at all. This chapter builds match decide, the K+1-armed generalisation, with one arm per listed option plus the mandatory uncertain arm. It asks what each arm means when scores compete on wording rather than meaning.

The question it must answer first is whether match is anything more than a switch statement with extra steps. The prompt allows exactly that verdict, and the chapter is built to be able to reach it. What follows is the semantics that would earn the construct its keep: exhaustiveness over listed options, ties that abstain instead of guessing, and stability measurements under paraphrase, order and distractors. If the measurements show no stability difference anywhere, the library form wins and match is a switch. That door stays open until JEV-18-01 runs.

What the scoring literature warns

Four papers fix what “matching on scores” can possibly mean. Each qualifies the one before it, and the last challenges the listing itself.

Surface-form competition: mass splits across synonyms

DOCUMENTED: Holtzman and colleagues name surface-form competition. Probability mass is finite, so synonymous surfaces (“computer” and “PC”, “USA” and “U.S. of A.”) split the mass that belongs to one meaning, and the argmax can fall to a worse meaning with a luckier surface. Their Domain Conditional PMI reweighs each option by its prior likelihood in the task domain, and it beats raw, normalised and calibrated scoring across GPT-2 and GPT-3 on more than a dozen multiple-choice datasets. [Holtzman et al., arXiv:2104.08315, Abstract, §§1–3.1, 4.3–4.4/§5, §§7–8]

The control experiment is the elegant half. On COPA Flipped, where competition is removed by construction (score the fixed continuation under varying premises instead of varying continuations under a fixed premise), all methods tie. That is proof that the gap was competition, not capability. Three datasets favour other methods, and the authors say so rather than smoothing it over.

For match semantics the lesson is structural: the listed options are surfaces, and matching on surfaces without accounting for competition measures familiarity as much as fit. Our perturbation battery is, in this light, a poor man’s COPA-flip. The paraphrase and distractor arms vary the continuations while the premise stays fixed, which is exactly the condition under which competition bites.

Order: the same samples, a different permutation

DOCUMENTED: Lu and colleagues show that order can swing few-shot classification from near-state-of-the-art to random chance, with the same samples and a different permutation. Larger models and more samples do not fix the variance. Good permutations do not transfer across models: a permutation scoring 88.7% on GPT2-XL drops to 51.6% on GPT2-Large, and the 175B/2.7B rank correlation is 0.05. [Lu et al., arXiv:2104.08786, Abstract, §§1–2, §5]

Their entropy probing buys a 13% relative improvement, but calibration, while raising accuracy, leaves variance high. Their error analysis adds the detail our epsilon rule needs: failing prompts mostly produce highly unbalanced predicted label distributions. Calibration repairs the balance without repairing the variance. That is why our near-tie epsilon must come from measured error rather than from a calibration map that looks healthy.

Two consequences follow. First, option order is a treatment, and any match experiment that does not perturb it is measuring one draw from 24 (or 77!) orderings. Second, fixes do not transfer across models, so stability must be tested per provider and never assumed from the literature.

Scoring in the other direction: channel models

DOCUMENTED: Min and colleagues reverse the scoring direction. Channel models compute P(input | label) rather than P(label | input), forcing the model to explain every word of the input. Across eleven text-classification datasets, channel models beat direct ones by 3.1 points on average and 7.2 in the worst case, and with prompt tuning by 13.3 and 23.5, through lower variance rather than higher peaks. Channel prompt tuning wins exactly where decisions live: imbalanced labels and unseen-label generalisation. [Min et al., arXiv:2108.04106, Abstract, §§1–3, §6.1, §7 + Limitations]

Two asides matter for our design. First, direct models with head tuning are surprisingly effective, often beating other direct variants, so the JEV-18-01 baseline set keeps direct scoring honest instead of straw-manning it. Second, the theory (Ng and Jordan) says channel models approach asymptotic error faster, which is why the stability half of P4 is a variance claim before it is a cost claim: more passes must buy less variance, or they buy nothing.

Their stated limitation is ours to inherit: non-classification priors are hard, which is why the noisy-channel arm of JEV-18-01 is a comparison and not a default. A scoring direction that needs a prior the task cannot supply is a method looking for a problem.

Paraphrase: the challenge to the listing itself

DOCUMENTED, and it challenges the listing itself: Janeiro and colleagues paraphrase only the correct answer and watch ARC-Easy accuracy rise 8 to 14 points; they paraphrase only the distractors and it falls 6 to 13. Same knowledge, different scores, at 70B and 120B scale as well as 1 to 8B. Their ParaEval remedy scores each option by its best paraphrase, so meanings compete with their best surfaces. [Janeiro et al., arXiv:2606.10657, Abstract, §§4.3, 5.1–5.2]

Two details sharpen the challenge. First, they use the maximum over paraphrases, not the average. Max asks whether the model knows the fact in any form; average would ask whether it knows it robustly in every form. Knowledge and robustness are measured separately or not at all. Second, their appendix warns that paraphrasing only the correct answer inflates scores by 5 to 9 points, which is why ParaEval paraphrases distractors too. Any future best-paraphrase arm of ours must do the same, or it will manufacture the stability it claims to measure.

The challenge to this chapter is direct. Rotation averaging and noisy-channel scoring both still match on author-chosen wordings. If the instability lives in the listing rather than the scoring, our whole perturbation battery measures the wrong thing, and ParaEval’s best-paraphrase rule is the unpreregistered rival to test next.

Wrong: “The match arm with the highest score is the meaning the provider endorsed.”

Correct: “The match arm with the highest score is the surface the provider scored highest. Meaning is what survives paraphrase, order and distractors, and that survival is measured, not assumed.”

    flowchart TD
    A[outcome] --> B{Decision?}
    B -- no --> C[uncertain arm]
    B -- yes --> D{value listed?}
    D -- no --> C
    D -- yes --> E{tied within epsilon?}
    E -- yes --> F[Abstain conflicting-evidence]
    E -- no --> G[run value arm]
  

What the earlier chapters already fixed

  • The First Token measured the readout side: bare-digit collisions, the bracketed uniform collision, and letter readout being impossible at 77 options.
  • Zero-Shot Decisions measured wording ranges of 13 to 29 points with the naive wording winning. Scores move with surfaces on our own providers, before any match layer exists.
  • if decide gave mandatory arms and the CRC selection rule; match generalises its two arms to K+1.
  • The Decision Expression gave Outcome, == versus same_choice, and the no-parser rule.
  • Transfer Across Decisions is the reason the noisy-channel comparison stays a comparison: no tuned models exist yet to test scoring directions on.

Chapter 18 adds the arm structure, the tie rules and the stability battery.

The build

The implementation is src/arbiter/match.py, tested by 8 tests in tests/match/ on hand-supplied numbers. It carries three semantics:

  1. Exhaustiveness. A Match has one arm per listed option plus on_uncertain, and is checked when it is built.
  2. Ties that abstain. tied_winners, resolve_ties, and the tie_epsilon declared on the Match itself, which match_outcome applies on the real path — one knob, no side channels.
  3. The decision table. outcome_kinds() names every row the runtime must handle.

Everything below is examples/ch18-semantic-match/walkthrough_ch18.py, which you can run as it stands. It reuses Chapter 16’s fake providers, so the distributions are hand-supplied and the output shows the semantics, not any real provider.

The setup

from arbiter.control import MissingArmError
from arbiter.expr import decide
from arbiter.match import Match, MatchError, match_decide, match_outcome, outcome_kinds, resolve_ties
from walkthrough_ch16 import FakeProvider, RefusingProvider  # Chapter 16's fakes

OPTIONS = ("refund", "exchange")
ARMS = (("refund", lambda d: "pay"), ("exchange", lambda d: "convert"))


def human(outcome) -> str:
    return f"human ({type(outcome).__name__})"

1. Exhaustiveness

A Match that forgets an option arm, duplicates one, or omits the uncertain arm is a construction error (MatchError, MissingArmError), raised before any outcome exists.

    match = Match(arms=ARMS, on_uncertain=human, options=OPTIONS)

    # 1. A match cannot be built unless it is exhaustive.
    print("1. exhaustiveness is checked at construction")
    for label, build in (
        ("forgot the exchange arm", lambda: Match(arms=ARMS[:1], on_uncertain=human, options=OPTIONS)),
        ("listed refund twice", lambda: Match(arms=(ARMS[0], ARMS[0]), on_uncertain=human, options=OPTIONS)),
        ("forgot the uncertain arm", lambda: Match(arms=ARMS, on_uncertain=None, options=OPTIONS)),
    ):
        try:
            build()
        except (MatchError, MissingArmError) as exc:
            print(f"   {label:<25} -> {type(exc).__name__}")

2. Dispatch, and what “unlisted” means

A Decision whose value has an arm runs it. A Decision for an unlisted value takes the uncertain arm: the provider answered, but the match did not ask. That rule matters more than it looks. Providers and match statements are written by different people at different times, and an option added to the provider but not to the match must not fall through to an arbitrary arm. It falls to uncertain, loudly. The alternative, dispatching on whatever the provider returned, arms or no arms, is precisely the plain-function behaviour Chapter 16 indicted: caller and provider coupled by an undocumented agreement about shape.

    # 2. Dispatch: listed winners run their arm; everything else is uncertain.
    print("2. dispatch")
    choices = ["refund", "exchange", "other"]
    cases = (
        ("listed winner (refund .70)", FakeProvider(choices, [0.70, 0.20, 0.10], "p1"), {}),
        ("unlisted winner (other .60)", FakeProvider(choices, [0.20, 0.20, 0.60], "p2"), {}),
        ("weak winner, threshold .5", FakeProvider(choices, [0.40, 0.35, 0.25], "p3"), {"threshold": 0.5}),
        ("provider refuses", RefusingProvider(), {}),
    )
    for label, provider, kw in cases:
        out = decide("What does the customer want?", "charged twice", choices, provider, **kw)
        print(f"   {label:<28} -> {match_outcome(out, match)}")

3. The decision table

Prompt task 5, done as code and not as a document: a table nobody executes is a wish. outcome_kinds() names six rows, and the example dispatches one of each without raising.

    # 3. The decision table: every kind of outcome has a row, and none can raise.
    print("3. the decision table")
    nota = ["refund", "exchange", "NONE_OF_THE_ABOVE"]
    table = {
        "Decision": decide("Q?", "s", list(OPTIONS), FakeProvider(OPTIONS, [0.7, 0.3], "t1")),
        "Refusal": decide("Q?", "s", list(OPTIONS), RefusingProvider()),
        "Abstain": decide("Q?", "s", list(OPTIONS), FakeProvider(OPTIONS, [0.6, 0.4], "t2"), threshold=0.9),
        "NoneOfTheAbove": decide("Q?", "s", nota, FakeProvider(nota, [0.2, 0.1, 0.7], "t3")),
        "Unknown": decide("Q?", "s", list(OPTIONS), FakeProvider(OPTIONS, [0.7, 0.3], "t4"), missing=("order_id",)),
        "Tie": resolve_ties({"refund": 0.5, "exchange": 0.5}),
    }
    assert sorted(table) == sorted(outcome_kinds())
    for kind, outcome in table.items():
        print(f"   {kind:<15} ({type(outcome).__name__:<14}) -> {match_outcome(outcome, match)}")

Note that “Tie” is a row but not a type. A tie is an Abstain whose reason is CONFLICTING_EVIDENCE, so the six rows are five types.

4. Ties, and why the rule has to be in the path

tied_winners(distribution, epsilon) returns the labels within epsilon of the maximum, best first. resolve_ties returns the single winner or Abstain(CONFLICTING_EVIDENCE). Epsilon defaults to 0.0, meaning exact ties only, and any larger value must come from measured error, never from convenience. “Near” is therefore a number with a provenance: within the error of the scores being compared, the runtime refuses to choose. The Lu error analysis is why this points at calibration error specifically, and why a healthy-looking ECE can still hide near-ties that are really ties. Epsilon must be measured on the same provider and split whose scores it will judge; an epsilon imported from another model’s calibration is a guess with a formula.

The first version of this module had resolve_ties as a helper beside the path and not in it. It was correct and tested in isolation, and the chapter’s flowchart said a tie reaches the uncertain arm. But decide() returns a Decision for the argmax even when two options share the top score, and match_outcome dispatched whatever it was given. An exact 0.5 / 0.5 distribution therefore ran an arm chosen by an order nobody asked for. Reviewing the chapter found the gap; the first fix was a match_decide composer with its own epsilon knob. Reading that fix back revealed a second smell: two knobs for one notion of “near”. The final shape declares tie_epsilon once on the Match statement and applies it inside match_outcome itself, so decide, match_outcome, and match_decide all take the same path. A helper that exists is not a rule that runs — and a rule that runs in two places with two knobs is a disagreement waiting for a caller.

    # 4. Ties: the rule is in the path, on a declared epsilon.
    print("4. ties, end to end (one knob: match.tie_epsilon)")
    tie = FakeProvider(OPTIONS, [0.50, 0.50], "tie-1")
    print(f"   exact tie, tie_epsilon=0.0 -> {match_decide('Q?', 's', list(OPTIONS), tie, match)}")
    close = FakeProvider(OPTIONS, [0.52, 0.48], "close-1")
    for eps in (0.0, 0.03, 0.05):
        near = Match(arms=ARMS, on_uncertain=human, options=OPTIONS, tie_epsilon=eps)
        print(f"   .52 / .48, tie_epsilon={eps:<4}      -> {match_decide('Q?', 's', list(OPTIONS), close, near)}")

5. Why scores need this care

The surface-form mechanism, on supplied numbers: meaning X at 0.45 beats meaning Y split as 0.30 and 0.28, but consolidate Y’s surfaces and Y wins at 0.58. Same meanings, different winner. The match layer sees only the listed options either way, which is exactly why the stability battery must perturb the listing.

    # 5. Why scores need care: meaning split across surfaces.
    print("5. surface-form competition (ILLUSTRATIVE numbers)")
    direct = {"X": 0.45, "Y-a": 0.30, "Y-b": 0.28}
    consolidated = {"X": 0.45, "Y": 0.30 + 0.28}
    print(f"   scored per surface   -> winner {max(direct, key=direct.get)}")
    print(f"   scored per meaning   -> winner {max(consolidated, key=consolidated.get)}")

What it prints

1. exhaustiveness is checked at construction
   forgot the exchange arm   -> MatchError
   listed refund twice       -> MatchError
   forgot the uncertain arm  -> MissingArmError
2. dispatch
   listed winner (refund .70)   -> pay
   unlisted winner (other .60)  -> human (Decision)
   weak winner, threshold .5    -> human (Abstain)
   provider refuses             -> human (Refusal)
3. the decision table
   Decision        (Decision      ) -> pay
   Refusal         (Refusal       ) -> human (Refusal)
   Abstain         (Abstain       ) -> human (Abstain)
   NoneOfTheAbove  (NoneOfTheAbove) -> human (NoneOfTheAbove)
   Unknown         (Unknown       ) -> human (Unknown)
   Tie             (Abstain       ) -> human (Abstain)
4. ties, end to end (one knob: match.tie_epsilon)
   exact tie, tie_epsilon=0.0 -> human (Abstain)
   .52 / .48, tie_epsilon=0.0       -> pay
   .52 / .48, tie_epsilon=0.03      -> pay
   .52 / .48, tie_epsilon=0.05      -> human (Abstain)
5. surface-form competition (ILLUSTRATIVE numbers)
   scored per surface   -> winner X
   scored per meaning   -> winner Y

Read it section by section.

  • Section 1: all three malformed matches fail at construction, with the two different error types.
  • Section 2: a listed winner runs its arm. An unlisted winner (other at 0.60), a weak winner and a refusing provider all reach the uncertain arm, each labelled by what it was.
  • Section 3: every row of the table is handled. The Tie row arrives as an Abstain.
  • Section 4: the exact tie abstains through every entry point — decide plus match_outcome and match_decide take the same path now. For a 0.52 / 0.48 split, tie_epsilon 0 and 0.03 let the winner through and 0.05 refuses to choose, because the gap of 0.04 is inside the stated error. One knob, declared on the match, applied on every dispatch.
  • Section 5: the split-surface and consolidated winners differ.

The experiment, designed but not run

JEV-18-01 is preregistered with status: NOT_RUN; the full block is in metadata/18-chapter.yaml. The run will call match_decide.

It uses the same items with three perturbations: one option reworded (the Chapter 5 w1 wording), option order reversed, and one distractor appended. A flip is an argmax that differs from the unperturbed decision.

There are three baselines: direct scoring, rotation-averaged direct scoring, and the noisy-channel direction. They are not three contestants but three questions. Direct scoring asks how bad the raw readout is. Rotation averaging asks whether the cheapest known fix, K rotations and no labels, removes it. Noisy-channel asks whether the scoring direction was the problem all along. If rotation fixes order flips but not paraphrase flips, the battery has localised two different instabilities to two different mechanisms, which is worth more than any single headline number.

Metrics are flip rate per perturbation, accuracy and latency. Acceptance is a flip rate reported per perturbation, with no aggregate that hides which perturbation did the damage. An averaged flip rate across paraphrase, order and distractor would let a stable order arm launder an unstable wording arm, repeating the exact sin Chapter 4’s per-class rule was written to prevent.

Predictions (HYPOTHESIS):

  • P1: paraphrase flips measurably (more than 2% on intent items).
  • P2: order flips measurably (more than 2%).
  • P3: distractors flip previously decided items (more than 1%).
  • P4: noisy-channel flips less but costs more forward passes per item.

Refutation: P1 to P3 fail on zero flip rates everywhere (match is stable; the battery found nothing). P4 fails if noisy-channel flips as much while costing more. And the prompt’s allowed verdict stands: if match reduces to a switch with no stability difference anywhere, recommend the library form. That recommendation would be a result about the construct, not a failure of the chapter.

PENDING_RUN: result for JEV-18-01, P1–P4 Command: python examples/ch18-semantic-match/run_ch18.py --perturb paraphrase,order,distractor --seeds 0,1,2 Fills: results/ch18.jsonl Stage plan (CPU-or-GPU per provider; perturbations on threshold items for selection, test evaluated once; each ≤15 min): S1 build perturbed requests (same items, three perturbations); S2 run direct/rotation/noisy-channel arms; S3 compute flip rates per perturbation with intervals; S4 accuracy/latency ledger; S5 write rows.

What would change your mind

Zero flips everywhere would collapse the chapter to its allowed verdict: match as a switch, the library form recommended, no construct earned.

Asymmetric flips, where paraphrase flips but order does not on our providers, would be more interesting than the prediction. They would localise the instability to the listing rather than the readout. That outcome has a named next step already: ParaEval’s best-paraphrase rule, currently unpreregistered, becomes the follow-up experiment, with the distractor-paraphrasing warning built into its design.

A noisy-channel win on stability at acceptable cost would promote the scoring direction from comparison arm to default, with Min’s prior-supply limitation as the documented boundary. It would also carry the honest admission that a direction needing P(input | label) per option is buying stability with passes, which is a cost claim Chapter 28 will audit.

No pivot is proposed. The software is demonstrated and the providers are pending.

Limitations

JEV-18-01 is NOT_RUN, and results/ch18.jsonl does not exist.

Required papers are PARTIAL by section; appendices, code and full ablations were not reproduced. ParaEval is read PARTIAL and stays unpreregistered. Invoking it as a conclusion would be borrowing an experiment we did not run, and invoking only its correct-answer paraphrasing without its distractor paraphrasing would repeat the optimistic bias its own appendix warns against.

Hand-supplied tests prove dispatch, tie and table semantics. They do not prove provider stability, wording effects or latency. The noisy-channel arm needs model runs by construction, since P(input | label) is not computable from a bare distribution, and its prior-supply limitation travels with it.

The epsilon-near-tie rule assumes a calibration-derived epsilon exists. Where it does not, only exact ties abstain, and the chapter does not pretend otherwise. The 2% / 2% / 1% flip floors are HYPOTHESIS thresholds with reasons (above noise, below literature effect sizes), not derived constants. Paraphrase and order get the higher floor because the literature moves double digits there, while distractor appends are the subtler intervention and get the lower one.

What the next chapters inherit

Chapter 19 inherits the open question the prompt leaves: which of match’s semantics are real (exhaustiveness, tie-abstention, stability-sensitivity) and which are a switch with extra steps. Only JEV-18-01 can grade that list, and the grading rubric is already written: stability differences per perturbation, or the library form. The rubric has an asymmetry worth noting. Exhaustiveness and tie-abstention are demonstrated in software regardless of the run, while stability-sensitivity lives or dies on provider measurements. A chapter that earns two of its three semantics and loses the third has still earned its keep, because the semantics were separated before they were tested. A switch with extra steps would have failed all three together.

Chapter 22 inherits the perturbation battery for where decide: filtering is matching with retrieval upstream, so the same flips apply, plus retrieval error, which Chapter 21 owns.

Chapter 28 inherits flip rates as a provider dimension beside accuracy and cost. A provider that is accurate but order-fragile loses differently from one that is merely inaccurate.

The closing question stays open by rule. The semantics are specified and tested; whether they are real is measured later, perturbation by perturbation, with no aggregate anywhere to hide behind.