← Jev From First Principles

if decide

Branch on a decision without lying: mandatory uncertain arms, the silent failure of boolean coercion, and thresholds with statistical backing.

Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.

The Decision Expression gave you a value that refuses to be ignored: an Outcome that is either a Decision or a typed non-answer. Now you have to use it, which means branching on it. And here is the failure this chapter exists to prevent, in the form every working programmer has written at least once:

if probs["refund"] > 0.5:
    pay_out()

That line looks honest. It is not. The 0.5 is not a fact about the world; it is a habit about numbers. Recalibrate the provider, say by fitting the temperature Chapter 10 taught you to fit, and the same items can cross the line in silence. Shift the distribution a little and they cross it again. No code changed, no test failed, no warning fired. The program now does something different for reasons recorded nowhere. A probability is not a boolean that has not finished dressing, and coercing one is the precise moment a program starts lying to itself.

This chapter builds the alternative: if decide, a branch with a mandatory uncertain arm and a threshold chosen for a stated risk, not a round number. The preregistered question, JEV-17-01, is whether mandatory handling changes program behaviour and whether the risk guarantee holds in practice. The hypothesis says yes to both halves: naive coercion fails silently under shift, and mandatory branches fix it at a coverage cost.

What the risk-control literature guarantees, and where it stops

Four papers give this chapter its threshold semantics. Three come from one school, each narrower and more honest than the last. The fourth shows what happens when the guarantee is asked to travel.

Conformal risk control: the correction term

DOCUMENTED: Angelopoulos and colleagues extend conformal prediction from miscoverage to any monotone bounded loss. Their rule selects λ̂ as the smallest conservativeness level whose corrected empirical risk, n/(n+1)·R̂ + B/(n+1), clears the target α, and it guarantees E[L(λ̂)] ≤ α, tight within O(1/n). [Angelopoulos et al., arXiv:2208.02814, Abstract, §§1–1.1, 4.1]

The correction term is the whole lesson in miniature: selecting on the same data you bound with costs you exactly 1/(n+1) of slack, no more and no less. Two boundaries come with the theorem. Under distribution shift the guarantee needs known likelihood ratios (their Propositions 2 and 3), which a deployer rarely has. And non-monotone risks are simply uncontrolled; the algorithm does not apply.

Learn then Test: the empty set is an answer

DOCUMENTED: Learn then Test reframes the same problem as multiple hypothesis testing. There is one null per candidate parameter value (R(λ) > α), one finite-sample p-value per null from a concentration inequality, and a family-wise error control such as Bonferroni across the grid. The output is a set of certified parameters, each an (α, δ)-risk-controlling prediction, and the set may be empty, in which case the procedure abstains from returning anything. [Angelopoulos et al., arXiv:2110.01052, Abstract, §§1–2.1, 3.2]

Their selective-classification worked example is our direct precedent: predict the argmax above λ, abstain (empty set) below, with error conditioned on predicting. They use five thousand calibration points, α = 0.15 and δ = 0.1 on ImageNet. The empty set is not a failure mode here; it is the procedure’s way of saying the evidence does not support any threshold.

Our MissingArmError and our (None, 0.0, None) return are the same idea at two levels. The branch refuses to be built without its uncertain arm, and the selector refuses to return a threshold the evidence does not support. In both cases the honest answer is a structured refusal, not a lucky number.

Conformal language modelling: when the output space cannot be listed

DOCUMENTED: Quach and colleagues carry the machinery to language models, where the output space cannot be enumerated. Their conformal sampling calibrates a stopping rule (keep sampling until the set probably contains an acceptable response) plus a rejection rule (prune low-quality candidates without breaking coverage). [Quach et al., arXiv:2306.10193, Abstract, §§1–4, Appendix A/B]

Three assumptions bound the result, stated in their appendix with unusual candour: i.i.d. data; an admission function that genuinely proxies quality; and a sample budget k_max large enough that the desired level is attainable, since otherwise the procedure returns null. Guarantees are probabilistic, never per-input. An if decide threshold inherits all three caveats the moment it touches a real provider.

The admission-function caveat maps directly onto decisions. Their guarantee needs a function that says whether a generation is “good enough”; ours needs the label to say whether a choice was right. A risk guarantee without a trustworthy correctness notion is a bound on a number nobody defined. Their k_max caveat maps onto our (None, 0.0, None): some risk levels are unattainable at a given calibration size, and the procedure must say null instead of inventing a threshold.

Group-wise guarantees: the one that contradicts the comfortable reading

DOCUMENTED, and it contradicts the comfortable reading: Salem and colleagues (read for Chapter 11, re-verified here) show marginal conformal risk control violating per-group budgets in up to 47% of trials under mild group-composition shift. Their fix, one threshold per hierarchy node with Bonferroni correction and leaf-first selection, restores the guarantee at a cost of 22 to 37 points of coverage on ARC-Challenge. [Salem et al., arXiv:2607.24562, Abstract, §1]

That coverage cost is the honest price of a guarantee that actually travels: narrower promises cost more abstentions. Our hypothesis predicts the same shape (mandatory branches fix silent failure at a coverage cost), and JEV-17-01 measures coverage beside risk for exactly that reason. A risk guarantee without a coverage column is an advertisement.

The guarantee is marginal over calibration-plus-test draws. It is not per-group, not per-input, and not portable across distributions. JEV-17-01 therefore tests the guarantee in-distribution and reports shift separately, because “the guarantee held” without a distribution qualifier would be precisely the lie this chapter exists to prevent.

Wrong: “Pick the threshold that hits your target error on the calibration set; it will hold wherever you deploy.”

Correct: “Pick the smallest cutoff whose corrected risk clears the target, on a split you never tune on, and report where exchangeability ends, because the guarantee ends there too.”

    flowchart TD
    A[calibration split scores] --> B{CRC bound <= alpha?}
    B -- yes, smallest tau --> C[pin threshold]
    B -- none --> D[accept nothing]
    C --> E[test split: branch]
    E --> F{Decision?}
    F -- yes --> G[value arm]
    F -- no --> H[uncertain arm]
    E --> I[shift split: report honestly]
  

What the earlier chapters already fixed

  • Abstain gave threshold arithmetic and the accounting insight that coverage is the price of risk.
  • The Decision Expression gave decide(), the total Outcome type, and docs/semantics.md as the compile target.
  • Confidence Is Not Probability gave the reason a naive 0.5 is arbitrary: the number is not a probability of being right, with or without calibration.
  • Calibration gave fitted maps plus the shift rows showing they expire off-distribution.

That last point is why the recalibration demonstration below uses two maps rather than one map plus noise. A single map can only show that thresholds matter. Two maps show that the branch outcome depends on which map the provider happens to wear while the code stays byte-identical. That is the precise sense in which the program lies to itself: not by computing wrongly, but by reporting the same code path as the same behaviour across two different decision policies.

Chapter 17 adds one mechanism on top: a branch that cannot drop the uncertain arm, and a threshold rule with a bound instead of a habit.

The build

The implementation is src/arbiter/control.py, tested by 6 tests in tests/control/ on hand-supplied numbers. It has three pieces:

  1. Mandatory arms. Branch and branch(), plus a context manager, decide_block, that is the shape if decide syntax will desugar to.
  2. The indicted baseline, kept on the premises. naive_coerce, so the failing example is runnable and not rhetorical.
  3. The CRC threshold rule. crc_threshold(scores, correct, alpha).

Everything below is examples/ch17-if-decide/walkthrough_ch17.py, which you can run as it stands. It reuses Chapter 16’s fake providers, so the distributions are hand-supplied and the output shows shapes, not any real provider.

The setup

from arbiter.control import (
    Branch,
    MissingArmError,
    branch,
    crc_threshold,
    decide_block,
    naive_coerce,
)
from arbiter.expr import decide
from walkthrough_ch16 import FakeProvider, RefusingProvider  # Chapter 16's fakes

CHOICES = ["refund", "exchange", "other"]
STATE = "my card was charged twice"


def ask(provider):
    return decide("What does the customer want?", STATE, CHOICES, provider, threshold=0.5)

1. Mandatory arms

A Branch carries on_value and on_uncertain, validated at construction. Omitting the uncertain arm raises MissingArmError before any outcome exists. The timing is the point: a runtime check fires after the program is deployed and the damage is scheduled, while a construction check fires in the editor, where the fix is one argument.

branch(outcome, handlers) routes a Decision to the value arm and every non-answer, whichever of the four kinds, to the uncertain arm.

    confident = FakeProvider(CHOICES, [0.70, 0.20, 0.10], model="confident-1")
    weak = FakeProvider(CHOICES, [0.40, 0.35, 0.25], model="weak-1")
    refuser = RefusingProvider()

    # 1. The branch cannot be built without its uncertain arm.
    print("1. mandatory arms")
    try:
        Branch(on_value=lambda d: f"pay out {d.value}", on_uncertain=None)
    except MissingArmError as exc:
        print(f"   missing arm  -> MissingArmError: {exc}")
    arms = Branch(
        on_value=lambda d: f"pay out {d.value} (score {d.score:.2f})",
        on_uncertain=lambda o: f"send to a human ({type(o).__name__})",
    )
    for provider, label in ((confident, "confident"), (weak, "weak"), (refuser, "refuser")):
        print(f"   {label:<9} -> {branch(ask(provider), arms)}")

2. The same thing as a context manager

decide_block validates on entry and yields a zero-argument dispatcher that runs exactly one branch. Arm errors propagate uncaught, because the context manager never suppresses an exception. This is the shape if decide will compile to: enter, check the arms, run one branch.

    # 2. The same thing as a context manager: the shape `if decide` desugars to.
    print("2. decide_block")
    with decide_block(ask(weak), on_value=lambda d: "act", on_uncertain=lambda o: "escalate") as run:
        print(f"   weak winner  -> {run()}")
    try:
        with decide_block(ask(weak), on_value=lambda d: "act"):
            pass
    except MissingArmError:
        print("   no uncertain arm -> MissingArmError on entry, before any outcome is read")

3. The villain, runnable

naive_coerce(prob, cutoff=0.5) exists so the failing example can be run. (The opening snippet writes >; the helper uses >=. The difference only matters at exactly 0.5.)

    # 3. The villain: coercing a probability to a boolean.
    raw = [0.62, 0.58, 0.51, 0.49, 0.30]
    recalibrated = [0.54, 0.52, 0.48, 0.46, 0.40]
    print("3. naive coercion, two calibration maps, identical code")
    print(f"   raw map          -> {[naive_coerce(p) for p in raw]}")
    print(f"   recalibrated map -> {[naive_coerce(p) for p in recalibrated]}")
    print("   item 3 flipped: no code change, no warning")

Item 3 flipped with no code change and no warning. That is P1 exhibited, not asserted: the same five items, two calibration maps, different program behaviour, silence throughout. Any recalibration that compresses probabilities toward 0.5 moves items across a fixed line, and a fitted temperature above 1 does exactly that.

4. A threshold from evidence

crc_threshold returns the smallest cutoff whose corrected empirical risk clears α, or (None, 0.0, None), meaning accept nothing, when no cutoff qualifies. The example builds a synthetic selection split of 40 items whose 11 errors cluster at low scores, then compares the corrected cutoff with the uncorrected one that simply picks the smallest cutoff whose raw empirical risk is at most α.

    # 4. Choosing the threshold from evidence instead of habit.
    n = 40
    scores = [round(0.50 + 0.0125 * i, 4) for i in range(n)]
    wrong = {0, 1, 2, 3, 5, 7, 9, 12, 15, 18, 24}  # errors cluster at low scores
    correct = [i not in wrong for i in range(n)]
    print(f"4. selecting a cutoff on a selection split (n={n}, {len(wrong)} errors, ILLUSTRATIVE)")
    print("   alpha   CRC cutoff  coverage  risk    | uncorrected cutoff  coverage  risk")
    for alpha in (0.20, 0.10, 0.05, 0.02):
        tau, cov, risk = crc_threshold(scores, correct, alpha)
        naive = next(
            (
                (t, sum(s >= t for s in scores) / n,
                 sum(1 for i, s in enumerate(scores) if s >= t and not correct[i])
                 / sum(s >= t for s in scores))
                for t in sorted(set(scores))
                if sum(1 for i, s in enumerate(scores) if s >= t and not correct[i])
                / sum(s >= t for s in scores) <= alpha
            ),
            None,
        )
        crc = f"{tau:.4f}  {cov:>7.2f}  {risk:.3f}" if tau is not None else "none      0.00    -    "
        unc = f"{naive[0]:.4f}  {naive[1]:>7.2f}  {naive[2]:.3f}" if naive else "none"
        print(f"   {alpha:<6}  {crc}   | {unc}")
    tiny_tau, _, _ = crc_threshold([0.9, 0.8, 0.7, 0.6, 0.5], [True] * 5, alpha=0.10)
    print(f"   n=5, every item correct, alpha=0.10 -> cutoff {tiny_tau} (slack 1/6 alone exceeds alpha)")

What it prints

1. mandatory arms
   missing arm  -> MissingArmError: if decide requires both on_value and on_uncertain arms
   confident -> pay out refund (score 0.70)
   weak      -> send to a human (Abstain)
   refuser   -> send to a human (Refusal)
2. decide_block
   weak winner  -> escalate
   no uncertain arm -> MissingArmError on entry, before any outcome is read
3. naive coercion, two calibration maps, identical code
   raw map          -> [True, True, True, False, False]
   recalibrated map -> [True, True, False, False, False]
   item 3 flipped: no code change, no warning
4. selecting a cutoff on a selection split (n=40, 11 errors, ILLUSTRATIVE)
   alpha   CRC cutoff  coverage  risk    | uncorrected cutoff  coverage  risk
   0.2     0.5750     0.85  0.176   | 0.5500     0.90  0.194
   0.1     0.7375     0.53  0.048   | 0.7000     0.60  0.083
   0.05    0.8125     0.38  0.000   | 0.7375     0.53  0.048
   0.02    none      0.00    -       | 0.8125     0.38  0.000
   n=5, every item correct, alpha=0.10 -> cutoff None (slack 1/6 alone exceeds alpha)

Read part 4 in three passes.

  • The corrected cutoff is always stricter. At every α the CRC cutoff is higher and its coverage lower. At α = 0.10, CRC accepts 53% of items at an empirical risk of 0.048; the uncorrected rule accepts 60% at 0.083. The difference in coverage is what the 1/(n+1) slack costs.
  • The two rules part ways at the strictest level. At α = 0.02 the uncorrected rule returns a cutoff, with empirical risk 0.000 on 15 accepted items, while CRC returns nothing. Zero observed errors among 15 items does not certify a 2% error rate, and the corrected rule knows it.
  • Small n certifies nothing. With only five items, all correct, α = 0.10 cannot be certified at all, because the slack alone is 1/6. The function says so by returning nothing rather than a lucky threshold.

Every number here is synthetic and the demonstration is about shapes. The point it makes is real, though: whether a threshold has been chosen from evidence or inherited from habit is a fact about the program, and it is a fact the branch can force the author to state.

A caveat the code states

Selective risk among accepted items is not monotone in the cutoff, while CRC’s guarantee needs a monotone loss. The finite-sample backing for cutoff selection therefore leans on the Learn-then-Test shape, per-cutoff binomial testing, for which Chapter 11’s bound already exists. The exchangeability this needs is exactly what JEV-17-01 must test. The module implements the point rule honestly and labels the guarantee as pending. That sentence is the difference between using the literature and wearing it.

The experiment, designed but not run

JEV-17-01 is preregistered with status: NOT_RUN; the full block is in metadata/17-chapter.yaml.

It compares three baselines on the same providers and splits: naive if p > .5, a naive tuned threshold (selected on the threshold split without correction), and the CRC threshold. The middle baseline matters more than it looks, and part 4 above is its toy version. It separates two claims that a two-way comparison would fuse. If the naive tuned threshold holds and the CRC one adds nothing, the lesson is “tune your threshold, skip the theory”: selection discipline, not statistical backing, did the work. If the naive tuned threshold fails where CRC holds, the correction term earned its keep and the 1/(n+1) slack is load-bearing rather than decorative.

Metrics are realised error against target, coverage, and behaviour change under recalibration and shift. Acceptance is in-distribution risk within the guarantee, with shift reported honestly either way.

Predictions (HYPOTHESIS):

  • P1: naive coercion flips silently. This is demonstrated on supplied distributions; the provider-level version is what the run checks.
  • P2: mandatory arms route every below-threshold outcome to the uncertain branch, and omission is a construction error.
  • P3: the CRC threshold realises selective risk within [α, α + 0.02] in-distribution. The tolerance acknowledges the O(1/n) tightness the CRC paper itself states; it is not hedging.

Refutation: P1 fails if naive coercion proves stable across recalibration maps; P2 on any outcome that bypasses the uncertain arm; P3 on in-distribution realised risk above α + 0.02.

The L2 verdict, whether enforced handling differs behaviourally from offered handling, is explicitly not decided here. These tests prove the enforced form exists and routes correctly; only real providers under shift can show that the difference matters.

PENDING_RUN: result for JEV-17-01, P1–P3 Command: python examples/ch17-if-decide/run_ch17.py --alpha 0.05,0.10 --seeds 0,1,2 Fills: results/ch17.jsonl Stage plan (CPU-only, reusing frozen Ch9 distributions + Ch10 maps; each ≤15 min): S1 assemble per-item scores per provider/seed; S2 select naive-tuned and CRC thresholds on the threshold split; S3 evaluate realised risk/coverage on test; S4 recalibration-flip and shift reporting; S5 write rows.

What would change your mind

A naive threshold that survives recalibration and shift would refute P1’s alarm. Zero point five would turn out to be robust in practice, and the chapter’s opening villain would retire. That outcome is not impossible: if providers are well calibrated and distributions stable, a fixed line barely moves anything, and the mandatory machinery buys little. The preregistration keeps that door open deliberately, because a hypothesis its author cannot lose is not a hypothesis.

A CRC threshold that misses in-distribution would refute P3 and point at the monotonicity caveat as the culprit, promoting the Learn-then-Test per-cutoff machinery from footnote to method.

And the result this chapter most wants, honestly: a run showing enforced and optional handling behave identically on real providers. That would not refute the branch, which routes correctly regardless. It would refute L2, collapse the case for if decide syntax, and leave the library form standing. The prompt permits exactly that verdict (“reject the construct if it is only an if with a try/except”), and the preregistration keeps the door open.

No pivot is proposed. The software is demonstrated and the providers are pending.

Limitations

JEV-17-01 is NOT_RUN, and results/ch17.jsonl does not exist.

Required papers are PARTIAL by section; proofs, appendices beyond those cited, and code were not reproduced. The CRC point rule here does not inherit CRC’s finite-sample guarantee for selective risk without the monotonicity and exchangeability conditions the chapter states. Salem’s per-group result bounds every guarantee claim to the calibrated distribution.

Fake-provider and hand-supplied-score tests prove routing and arithmetic, not provider behaviour; miscalibration, cost and shift belong to Chapters 4 to 10. The synthetic selection split in part 4 is a fixed pattern chosen to show shapes. It is not evidence about any provider.

The context manager is a prototype, not syntax: no parser and no static exhaustiveness checking. A None threshold (accept nothing) is honest but undeployed. No caller policy for total refusal is designed here, and a program that refuses everything is a denial of service with good statistics. The recalibration demonstration uses two hand-built maps; real recalibration drift has its own shape, which only the run can supply.

What the next chapters inherit

Chapter 18 inherits the branch contract generalised: match decide must be exhaustive over all five Outcome members, with the same construction-time error for a forgotten arm.

Chapter 19 inherits the open question of which outcome types earn primitives. if only ever needed two arms (value, uncertain), and whether the uncertain arm should split further is a types question.

Chapter 29 inherits the threshold-selection machinery for the router: risk-controlled cutoffs chosen on held-out splits, reported with coverage costs.

The prompt’s closing question, whether you would accept a language that let you omit the uncertain branch, is answered by construction here and by measurement later. The constructor already refuses, and JEV-17-01 will say whether the refusal was worth it.