← Jev From First Principles

The Decision Contract

A provider can answer, refuse, or violate the contract. Fake-model swaps test which distinctions the caller actually needs.

You replace the refund predicate with a classifier. The next ticket arrives, but your question names a label the classifier has never seen. Should the application guess, catch an exception, or pass the ticket to a human? A distribution over the wrong labels would look reassuring while answering a different question.

Stop Generating kept that application’s action separate from its provider. The Jev Provocation supplied a first request and response shape. The Smallest Decision and Zero-Shot Decisions then supplied unlike providers. The First Token exposed a further problem: some ways of presenting an option set cannot be read by a particular provider. You need to know whether it answered your request at all before asking how good its answer was.

This chapter tests the interface with deterministic fakes and recorded scores. OBSERVED means code executed here on CPU; it never means an NLI, embedding or language model was evaluated in this session. Hosted Jev is a wire-shape stub. Accuracy, calibration, model latency and live hosted behavior are NOT_OBSERVED.

The expectation and its limits

HYPOTHESIS — H-contract: a neutral contract needs an explicit way for a provider to decline a valid request, and presentation belongs to the provider. The committed prediction expected two or three providers to break the first draft. It also predicted that provider-chosen presentation would need fewer caller changes than caller-required presentation.

The separation has prior art. DOCUMENTED: DSPy signatures declare a task while modules and compilation determine how a model implements it. The book imports that distinction from DSPy From First Principles, rather than claiming it as a new mechanism. Khattab et al., §3.

Meyer supplies the vocabulary of caller obligations, supplier guarantees and invariants. His discussion also permits a broader interface that handles expected special cases. DOCUMENTED: contract theory does not uniquely require refusal to be a returned value; responsibility boundaries remain a design choice. The primary article was accessible in this resumption, correcting the earlier reading note. Meyer, pp. 42–44 and 49–51.

DOCUMENTED: JSONSchemaBench tests whether engines admit valid instances and reject invalid ones; both excessive restriction and insufficient restriction are failures. That is why our suite includes supported requests alongside impossible ones. Geng et al., §5.3.

The recent structured-output study by Chavan provides the strongest warning: schema-valid output can omit part of a requested task, and field-presence metrics can miss empty substance. Our assertions check consistency, not truth. Chavan, §§4.3–5.1.

All required papers were read from primary full-text copies, with the inspected sections recorded as PARTIAL for this resumption. The modern paper was read FULL_TEXT. The chapter’s source-verification companion records the sections, fetch failures and correction of an inherited misattribution to JSONSchemaBench.

What the contract promises

PROPOSED: retain the earlier question and answer names, and put a checked boundary around them. ContractProvider.decide returns an answer or Refusal under every requested question id. validate_result binds the answer to the request: its probability keys must cover exactly the declared options. Checking only that the returned distribution sums to one cannot detect omitted options.

Decision[State, Value] projects a choice into value, score, alternatives, provider, evidence and trace, alongside its state and declared answer set. Score is the selected probability. Alternatives retain the request’s label order. Evidence records provenance; trace identifies the checked request. Neither is a proof that the label means what you intended.

The outcomes stay distinct:

Situation Representation What the caller learns
Invalid question or blank state Precondition exception Repair the request
Supported choice Choice answer and distribution A declared option was selected
Provider cannot satisfy a valid request Typed refusal, without a distribution Route or record the inability
Explicit abstention option selected Abstention.ABSTAIN as a declared value An answer inside the declared space
Broken code or transport failure Exception An expected refusal must not conceal a defect

Refusal describes inability to supply the requested answer; abstention reserves an answer value. This chapter chooses their representation. Later chapters determine their policy and semantics. No calibration or open-set claim follows from the type.

Presentation is optional provider evidence, not a mandatory prompt field for every classifier. The option scorer reports its printed markers and explicitly refuses an impossible letter readout. A caller may impose a presentation through the experimental wrapper; the provider must honour it or refuse.

This maps directly to Language’s Representation Ladder:

Element Rung
State, instructions and option descriptions natural prose
Question ids and declared options structured requirements
Value, score, alternatives and refusal fields schemas and constraints
Running checks of distributions and request binding executable contracts
Provider, evidence and trace records machine-native operational state
Application action after the outcome state-transition representations

The contract is not a formal specification or proof of semantic correctness. planning/decision-contract.md explains the mapping and extension points.

The contract, running

The table above is easier to believe when each row is a call. Everything below is examples/ch08-the-decision-contract/walkthrough_ch08.py, which you can run as it stands. The providers are deterministic fakes that return whatever they are told to, so the example shows which distinctions the contract draws and says nothing about any real provider.

from arbiter.contract import (
    Abstention,
    ChoiceAnswer,
    ChoiceQuestion,
    DecisionRequest,
    DecisionResult,
    Refusal,
    RefusalReason,
    choice_confidence,
    decision_for,
    validate_result,
)

CRITERIA = {"billing": "Charges and payments", "returns": "Wrong or damaged items"}


def request_for(criteria, state="my card was charged twice"):
    return DecisionRequest(
        state=state,
        questions={"route": ChoiceQuestion(id="route", instructions="Which team handles this?", criteria=criteria)},
    )


def answer_with(probabilities, choice=None):
    """A provider's answer. `choice` defaults to the argmax, as the contract requires."""
    choice = choice or max(probabilities, key=probabilities.get)
    return DecisionResult(
        model="fake-1",
        answers={"route": ChoiceAnswer(choice, probabilities, choice_confidence(list(probabilities.values())))},
    )


def attempt(label, build):
    try:
        build()
        print(f"   {label:<34} -> accepted")
    except ValueError as exc:
        print(f"   {label:<34} -> ValueError: {exc}")


def main() -> None:
    request = request_for(CRITERIA)

    # 1. A bad request is the caller's bug, so it raises before any provider is asked.
    print("1. an invalid request is a precondition error")
    attempt("blank state", lambda: request_for(CRITERIA, state="   "))
    attempt("only one option", lambda: request_for({"billing": "Charges"}))

    # 2. A supported choice: the answer is bound to THIS request and projected to a Decision.
    print("2. a supported choice")
    result = validate_result(request, answer_with({"billing": 0.70, "returns": 0.30}))
    decision = decision_for(request, result, "route")
    print(f"   value={decision.value!r} score={decision.score:.2f} provider={decision.provider!r}")
    print(f"   alternatives={decision.alternatives}")
    print(f"   evidence={decision.evidence[0]!r}")

    # 3. A refusal is a typed answer without a distribution, not an exception.
    print("3. the provider cannot satisfy a valid request")
    refusal = Refusal(reason=RefusalReason.LABELS_UNSEEN, message="never fitted on these labels")
    refused = decision_for(request, validate_result(request, DecisionResult("fake-1", {"route": refusal})), "route")
    print(f"   value is a Refusal: {isinstance(refused.value, Refusal)}, score={refused.score}, alternatives={refused.alternatives}")

    # 4. Abstention is a value inside the declared answer space, which is a different thing.
    print("4. abstention is a declared option, not a refusal")
    with_abstain = {**CRITERIA, Abstention.ABSTAIN.value: "None of these fits"}
    request_a = request_for(with_abstain)
    picked = answer_with({"billing": 0.20, "returns": 0.10, "ABSTAIN": 0.70})
    abstained = decision_for(request_a, validate_result(request_a, picked), "route")
    print(f"   value={abstained.value!r} (in the answer set: {abstained.value in abstained.answer_set})")

    # 5. A provider that breaks the contract is a defect, and it is caught, not smoothed over.
    print("5. contract violations")
    attempt("probabilities sum to 1.2", lambda: answer_with({"billing": 0.70, "returns": 0.50}))
    attempt("winner is not the argmax", lambda: answer_with({"billing": 0.70, "returns": 0.30}, choice="returns"))
    attempt("answers a different option set", lambda: validate_result(request, answer_with({"billing": 0.6, "shipping": 0.4})))
    attempt("silently drops an option", lambda: validate_result(request_a, answer_with({"billing": 0.7, "returns": 0.3})))
1. an invalid request is a precondition error
   blank state                        -> ValueError: state must not be empty: a decision about nothing is nothing
   only one option                    -> ValueError: choice needs 2-255 options, got 1
2. a supported choice
   value='billing' score=0.70 provider='fake-1'
   alternatives=(('billing', 0.7), ('returns', 0.3))
   evidence='provider-returned distribution; not a truth guarantee'
3. the provider cannot satisfy a valid request
   value is a Refusal: True, score=None, alternatives=()
4. abstention is a declared option, not a refusal
   value='ABSTAIN' (in the answer set: True)
5. contract violations
   probabilities sum to 1.2           -> ValueError: choice: probabilities sum to 1.2, not 1.0
   winner is not the argmax           -> ValueError: choice 'returns' is not the most probable option ('billing' is)
   answers a different option set     -> ValueError: answer must cover exactly the declared option set
   silently drops an option           -> ValueError: answer must cover exactly the declared option set

Read it against the table.

  1. A bad request is the caller’s bug. A blank state or a one-option question raises before any provider is asked. Nothing was refused, because nothing was asked.
  2. A supported choice comes back as a Decision with its score, its alternatives in the request’s own order, the provider that answered, and an evidence note that says outright it is not a truth guarantee.
  3. A refusal is an answer without a distribution. The Decision carries the Refusal as its value, with no score and no alternatives. The caller learns the provider could not answer, and nothing in the result resembles a probability.
  4. Abstention is a value inside the declared answer set. Here the request lists ABSTAIN as one more option and the provider picks it. That is an answer, and a different thing from the provider declining to answer at all. This chapter fixes how the two are represented; later chapters decide when to use them.
  5. Violations are defects, not outcomes. Probabilities that sum to 1.2, a winner that is not the argmax, an answer over the wrong option set, and an answer that silently drops a declared option all raise. The last two are the reason validate_result binds the answer to this request: a distribution that sums to one cannot reveal that it answers a different question.

What actually broke

OBSERVED — JEV-08-01: all eight existing providers at ca1e676 raised an exception for a valid request with unsupported labels. All could answer their supported fixture, but none expressed refusal through the first-draft result. The hosted stub did not exist at that baseline and is excluded from this count. The prediction of two or three failures was refuted, not softened to match.

The completed suite passed 49 tests. It caught four of four seeded defects: invalid probability mass, an out-of-set winner, wrongly attached label scores, and silent truncation. The truncating provider returned a normalized distribution over only part of the declared set; exact request binding caught it. The mutated order case needed the known-score oracle. Structural invariants alone accepted that internally consistent but semantically wrong answer.

The loose schema control accepted three malformed outcomes that the checked contract rejected. This demonstrates the value of these additional checks against that control. A stronger request-specific schema could reject some of them too; the result does not prove that a new language construct is necessary.

These numbers come from results/ch08.jsonl: first_draft_refusal_failures, property_tests_passed, seeded_bugs_caught, and schema_control_contract_failures.

The swap, including the caller’s cost

OBSERVED — JEV-08-02: every provider saw the same 200 BANKING77 threshold texts. Its score computation was faked; the frozen train subset was selected but not used to estimate model performance. The rule provider and recorded stub refused the answer space. Refusing counts as an explicit outcome, not success at classification.

Provider Adapter lines Edits per subsequent swap P1 refusals P2 refusals
Majority 23 0 0 200
Lookup 26 0 0 200
Rule 36 0 200 200
TF-IDF + LR 20 0 0 200
fastText-style 58 0 0 200
Embedding similarity 46 0 0 200
NLI 49 0 0 200
Option scoring 39 0 0 0
Hosted Jev stub 11 0 200 200

Sources: adapter_lines_corrected, adapter_lines_by_arm, and caller_changes_per_swap, including each row’s arm and refusals fields. The adapter rule includes provider-specific fit bodies, rather than counting only the final return statement. Its original absolute size predictions mostly failed. An audit-counter repair preserved the original rows and appended corrections; the dated amendment explains the missed helper and model-loading exclusions.

P1 lets the provider choose presentation. P2 requires bare-number markers through a shared wrapper costing 11 additional lines. Both arms use the same caller with zero edits per swap. The strict predicted caller-change advantage for P1 was refuted by a tie. P2’s coverage loss is conditional on that fixed constraint; it does not establish that all caller constraints are harmful.

The literal older runner could not consume a refusal: it reads answer.choice unconditionally. Migrating its loop required five added nonblank lines under the declared diff measurement. After that migration, swapping providers required no further edits. The unchanged Chapter 1 application accepted the common refund-outcome adapter, including escalation on refusal. Literal zero migration changes was therefore refuted; subsequent interchangeability was observed.

Provider knowledge still reaches the original benchmark caller:

Location in run_ch04.py Finding Treatment
labels_of, line 64, and its call sites Caller discovers the answer space through provider.labels Provider detail leaks into request construction
build, lines 114–122 Constructs providers by name Recorded, excluded as factory setup
Seed branches, lines 143, 215, 310 Runner knows which providers are stochastic Setup knowledge remains in the runner
except ValueError, lines 160, 224, 318 Generic fit/budget failures become provider_refused Refusal and failure are conflated
decide_all, line 57 Reads a success-only field Refusal is not handled

The raw AST scan reported 12 candidates, including five excluded factory branches; removing those leaves seven. Manual reading counted 14 leak locations, including indirect label-helper calls and unhandled refusal. The new common caller has zero AST and manual leaks under these rules. The original runner remains unchanged, so its leaks remain. This also corrects the brief’s description: its handlers catch generic ValueError; they do not match an LR-only error string. See the ast_leaks and manual-audit rows for exact locations.

Regression evidence and its boundary

OBSERVED — JEV-08-03: recorded-score replay preserved 1,540 decisions across 11 task/provider combinations, with probability differences below the preregistered tolerance. These are the golden_decisions_identical and golden_max_probability_delta rows. Replay checks label binding, normalization and winner selection. Because it injects recorded scores, it cannot detect a changed encoder or training algorithm. The earlier end-to-end golden capture and check belong to commits 9296ae4 and 303ddc5; this session does not claim to have repeated them.

The five earlier self-test files passed without edits. Chapter 5’s encoder was replaced at its loading seam with a fake, so its pretrained-model smoke remains NOT_OBSERVED here. The compatibility shim measures 13 physical code lines; legacy imports, successful outputs and unsupported-label exceptions survive. NLI’s committed code was wrapped and fake-tested without editing its Chapter 6 file in this resumption. Earlier results and preregistrations were not changed.

The stub parses a transcribed response shape from evidence/jev-access.md using the Chapter 2 parser. Its recorded choice/noul shape fits; no answer-shape change is needed. It does not establish byte equality with a raw HTTP payload, which that note did not retain. No Jev call occurred. TypeSafe’s inspected successful-response types have no refusal arm, but that does not establish that the server cannot decline through an HTTP error. Vendor SDK types.

The decision to carry forward

PROPOSED verdict: H-contract is PARTIALLY_SUPPORTED. Explicit inability is useful under the broad neutral interface. Refusal as a value is our representation choice; a shared typed exception remains an untested alternative. The swap does not establish exclusive provider ownership of presentation. Keep presentation as optional audit evidence, with feasibility checked by the provider.

I recommend accepting part (a) of Chapter 7’s pivot narrowly: make presentation feasibility visible in the chapter and refuse an impossible readout. Defer a stronger rule that presentation must be a mandatory neutral-contract field or must never be constrained by a caller. The author still decides; the proposal’s status has not been changed to accepted.

You can inspect the example and verify the recorded experiment without a model:

python examples/ch08-the-decision-contract/demo_ch08.py
python examples/ch08-the-decision-contract/verify_ch08.py

The durable artifacts are the checked contract, common adapter, property suite, golden replay, caller audit and results record. The failed fixture assumption about integer label keys is recorded, as are the presentation-arm constraint, partial paper reading and the limits of fake-model regression. The chapter stays a draft for author review. Confidence Is Not Probability now needs to distinguish a checked score from a calibrated one.

What is now permanently separated, and what could still leak?