← Jev From First Principles

The Decision Specification

What must a declarative decision specification say, so a validator can catch mistakes the library API cannot?

The Decision Specification

This chapter is model-free: a declarative DecisionSpec and a validator, with six seeded bug classes and a reported detection rate. No model runs; the numbers are the validator’s own outputs.

The problem

Every chapter so far asked a provider a decision through the library API. The API checks a single call in isolation. But a decision in a program has a life beyond one call: it may abstain, and the abstained request must go somewhere; it may fall back to another provider; its answer classes must each have a consequence; its thresholds must be coherent with its risk target. None of that is expressible in a single library call, and none of it is checked by the call’s own validation.

This chapter asks: what must a declarative decision specification say, so a validator can catch mistakes the library API cannot?

What we expect and why

Our starting hypothesis is that a typed declarative spec makes a declared class of authoring bugs detectable at authoring time. Three sources frame it:

  1. Cutler et al. (2024) — Cedar is an authorization policy language whose validator uses optional typing so policy writers “avoid mistakes, but not get in their way”, and whose design is analyzable. The exact prior art shape: a declarative policy plus a validator that runs before the policy does. (Read at abstract level.)

  2. Molino et al. (2019) — Ludwig’s declarative configuration files name the data types and the model, so an inexperienced user gets a working model without writing code. The type-based declarative config is the pattern. (Read at abstract level.)

  3. Mitchell et al. (2019) — Model Cards are short documents describing a model’s intended use, evaluation and caveats. The spec here is the decision-analogue: what the decision is for, and what it requires. (Read at abstract level.)

The build

src/arbiter/spec.py defines the declarative spec and the validator.

from arbiter.spec import AnswerClass, DecisionSpec, validate
    # 1. A clean spec passes.
    print("1. a clean spec")
    clean = DecisionSpec(
        name="policy", inputs=("article",),
        answer_classes=(AnswerClass("publish", "publish", "publish"),
                        AnswerClass("reject", "reject", "reject")),
        abstention=True, abstention_route="human",
        fallback_chain=("rule", "judge", "human"),
        keep_threshold=0.9, abstain_threshold=0.4, risk_target=0.1,
        evidence_required=True,
    )
    print(f"   issues = {validate(clean, frozenset({'rule','judge','human'}))}")
    # 2. The same spec, missing its abstention route: B1.
    print("2. missing abstention route")
    broken = DecisionSpec(clean.name, clean.inputs, clean.answer_classes,
                          True, None, clean.fallback_chain,
                          clean.keep_threshold, clean.abstain_threshold,
                          clean.risk_target, clean.evidence_required)
    print(f"   issues = {[i.code for i in validate(broken, frozenset({'rule','judge','human'}))]}")
1. a clean spec
   issues = []
2. missing abstention route
   issues = ['B1']

The bug classes

The validator rejects five classes that a single library call cannot see (B4 covers two sub-conditions, thresholds and risk target):

Code Bug Why the API can’t catch it
B1 missing abstention route the API sees one call, not where the abstained request goes
B2 unreachable fallback (repeated or unknown provider) the API does not know the program’s provider registry
B3 overlapping answer sets the API checks shape, not whether two options mean the same thing
B4 inconsistent thresholds / risk outside (0,1] the API checks each value, not their joint coherence
B5 answer class with no handler the API returns the answer; it does not know the caller must handle it

The run

run_ch32.py seeds one spec per bug class and a clean spec, and reports detection (results/ch32.jsonl):

bug class detected codes found
B1 yes [B1]
B2 yes [B2]
B3 yes [B3]
B4 yes [B4]
B5 yes [B5]
clean accepted []

Detection rate 5/5; the clean spec is accepted (no false positive). All six predictions hold — the validator catches each seeded bug and nothing it should not.

The place on the Representation Ladder

The Language book names the ladder plainly: “There is a ladder of representations between ordinary language and executable machine state: prose, controlled prose, structured requirements, constraints, executable contracts, state-transition models, formal specifications and operational state.”

A DecisionSpec sits between structured requirements and constraints — rungs where expressiveness is still high and checkability is rising but the representation is not yet executable. The validator is what moves it toward executable: like Cedar’s validator, it makes the spec analyzable before it is run. The book’s warning applies to it: “Making a statement easier to check can make it worse at preserving what was meant, so intent preservation, checkability and conformance must be kept apart.” The spec trades those — a tight schema can reject a spec whose meaning is subtle — and the chapter keeps the three separate rather than blaming a rejection on a meaning error.

Wrong / Correct. Wrong: “A decision specification is the executable program.” Correct: “A decision specification is a checkable representation of intent — Cedar’s validator for decisions. It sits between structured requirements and constraints on the ladder, it can catch bugs a library call cannot, and it is not, by itself, the program: the compiler (Chapter 33) is what makes it executable.”

The distinction this chapter keeps

Authoring-time vs runtime. The spec and validator catch mistakes before anything runs. The library API catches mistakes at the moment of one call. Cedar argues the former. This chapter does not test that argument: the five bugs were seeded by the same author who wrote the validator, so 5/5 shows the checks do what they were written to do, not that they find bugs people actually make. Zero false positives is on one clean spec.

Specification vs provider vs program. The spec names the decision’s requirements; the provider answers; the program (Chapters 25-26, 34) composes. A spec that is forced to contain program facts (the provider registry, the handlers) is how the validator gains its power, and that is exactly why the API alone could not have the same checks.

What to carry forward

Chapter 33 compiles the spec into a plan: compile(spec, profiles) -> plan, choosing among rule, classifier, retrieval, model, cascade and human from supplied cost and quality profiles. The validator hands it a strong guarantee: the input plan’s intent is already checked — the compiler’s job is the choice, not the meaning.

Close by

What is the minimum a decision must record to count as evidence? The spec answers it declaratively: the decision’s name, typed inputs and answers, its abstention route, its fallback chain, its required risk and thresholds, its evidence requirements — the fields the validator checks, and the ones Chapter 26’s graph records at run time.

Limitations

  • The validator is a hand-written rule set over the declared six bug classes; it is not a general static-analysis system (it will miss bugs outside the classes).
  • The three papers are read at abstract level; Cedar’s validator internals were not studied for the rule set.
  • The Representation Ladder placement uses the Language book’s words (11-chapter) read in the source; no claim is made beyond that placement.
  • No model ran; the 5/5 detection is a property of the six seeded specs, not of all possible specs.