Jev From First Principles
A build-and-investigate book about typed, uncertain decisions as a programming abstraction: reconstructing the Jev idea from first principles, testing it against classifiers, embeddings, NLI and language-model alternatives, and building a decision stack that keeps uncertainty, evidence and provider choice explicit.
Most software asks models to generate language even when the application needs something much smaller: a bounded answer that can drive code. Approve or reject. Relevant or irrelevant. Permit, deny or escalate. The prose may be useful to a person, but the program usually consumes a value.
Jev From First Principles starts with that mismatch.
Jev made the idea visible by treating model output as a typed, probabilistic decision rather than as prose. But this book does not begin by assuming that Jev is a new model category, that its implementation is superior, or even that semantic decisions deserve new programming syntax. Those are the questions under investigation.
What does programming look like when semantic interpretation and uncertain decisions become first-class computational primitives?
The book answers that question by building, measuring and rejecting things. If a logistic-regression baseline wins, it stays. If an embedding model is enough, the more elaborate provider does not get to replace it merely because it is newer. If a proposed construct reduces cleanly to an ordinary library operation, the construct is rejected.
That last point matters. The book eventually proposes forms such as decide, if decide, match decide, where decide and for decide. They are not treated as language features because the syntax looks attractive. Each has to earn its place through semantics, implementation and evidence.
What the book builds
The concrete system is called arbiter: a tested library of typed-decision primitives and supporting runtime machinery. It grows chapter by chapter from one decision contract into calibration, abstention, semantic control flow, retrieval, decision graphs, provider routing, cascades, a declarative specification, a cost-based compiler and a full semantic program.
The architecture deliberately separates four things that ordinary model code often collapses:
decision specification what must be decided
provider how the answer is produced
runtime how uncertainty and control flow are handled
evidence what was observed and can be replayed
The same decision contract can therefore be answered by a rule, classifier, embedding model, NLI model, language model, cascade, or eventually a Jev-style decision model without changing the application-facing meaning of the decision.
The long-term idea is larger than Jev itself. A program should be able to ask for a semantic decision while remaining explicit about the output type, confidence requirements, abstention behaviour, provenance and fallback. The provider is an implementation choice, not the definition of the decision.
The questions the book keeps open
Four hypotheses concern the model:
H1 NOTHING NEW decision models reduce to familiar techniques plus packaging
H2 USEFUL ABSTRACTION the methods are familiar but the software abstraction is valuable
H3 GENERAL CAPABILITY decision-making transfers across otherwise different tasks
H4 COMPROMISE REGION decision models occupy a useful general + fast + cheap region
Three concern the language:
L1 SYNTAX ADDS NOTHING the proposed constructs are only library calls
L2 ENFORCE UNCERTAINTY uncertainty must be structurally handled, not merely documented
L3 A COMPILER CHOOSES provider selection belongs in compilation/planning, not caller code
The book is allowed to refute every one of them.
How the investigation works
Each chapter follows the same discipline:
question
-> prior work
-> pre-registered prediction
-> implementation or experiment
-> observed result
-> chapter verdict
-> durable artifact
Claims are kept separate by kind:
- Documented — supported by an external paper, specification, API or source the chapter names.
- Observed — produced by a recorded run in this repository.
- Proposed — a design, interpretation or language construct introduced by the book.
A green unit test proves the code around a model, not the model. A simulated cascade proves the cascade logic over supplied numbers, not the behaviour of a real provider. A design-only chapter remains design-only until the experiment runs. The book keeps those boundaries visible because collapsing them would make the central question impossible to answer honestly.
The nine parts
Part I — The Decision Problem (chapters 1 to 4).
Why applications often need bounded answers rather than prose; what Jev proposes; why the idea became controversial; and how much ordinary rules and classifiers already achieve.
Part II — From Classification to Semantic Decisions (5 to 8).
Zero-shot decisions, NLI, first-token readout and the decision contract that separates what is asked from how it is answered.
Part III — Uncertainty Is Part of the Type (9 to 12).
Why confidence is not probability; calibration; abstention; and the difference between unknown, none-of-the-above and refusing to decide.
Part IV — Tuning the Decider (13 to 15).
Hidden-state readout, tuning small arbiters across base models and the transfer experiment that would decide whether a general decision capability really exists.
Part V — Semantic Control Flow (16 to 19).
The decide expression, if decide, match decide and the minimum set of decision types. The emphasis is operational semantics: ties, thresholds, abstention and mandatory uncertain paths.
Part VI — Embeddings Meet Decisions (20 to 23).
A decision cannot read everything. Retrieval becomes candidate generation; embeddings and BM25 over-recall; a decision filter narrows the set; where decide is tested; for decide is rejected when ordinary library forms prove equivalent.
Part VII — Decisions Compose (24 to 27).
Error propagation, semantic pipelines, replayable decision graphs, provenance and the question of whether several decisions really share useful internal work.
Part VIII — Providers (28 to 31).
One contract, many mechanisms: where classifiers, embeddings, model judges and decision models would compete; then routers, cascades and learned provider selection.
Part IX — The Decision Language (32 to 35).
A declarative decision specification, a cost-based compiler over provider profiles, a complete semantic program and the final verdict over the evidence ledger.
What survived the build
The syntax did not all survive.
decide survived
if decide survived as enforced uncertain control flow
match decide survived with mandatory uncertainty and tie handling
where decide survived as a typed, abstaining semantic filter
for decide rejected; ordinary filter/takewhile/sorted forms were enough
That rejection is part of the result. The book is not trying to maximize the number of constructs it invents.
What was actually measured
The repository contains real CPU experiments and deterministic runtime tests, but it does not yet contain the decisive real decision-model wave.
Measured results include a tuned TF-IDF + logistic-regression baseline on a 77-way intent task, zero-shot embedding and NLI comparisons, calibration and conformal-coverage degradation under distribution shift, a SciFact retrieval-and-filtering pipeline, and deterministic tests of decision graphs, routers, cascades, the decision specification, compiler and semantic program.
Several important experiments remain deferred: the model sweeps around first-token decision readout, the hidden-state and transfer experiments, the shared-state work, and the final cross-provider winner table. No hosted Jev provider exists in the repository yet.
For that reason the final chapter keeps every book-level model hypothesis H1-H4 and language hypothesis L1-L3 at INSUFFICIENT_EVIDENCE. The book has built a serious typed-decision system and learned several things about uncertainty, retrieval, composition and control flow. It has not yet earned the stronger claim that a distinct decision-model family is necessary or that semantic decision syntax should become a programming-language primitive.
The narrowest defensible claim
On the evidence currently in the book, arbiter is best described as a tested library of typed-decision primitives: explicit answer contracts, mandatory uncertainty handling, abstention and calibration, replayable decision provenance, declarative specifications, provider-independent execution, a cost-based compiler over supplied profiles, and a semantic runtime boundary that can permit, deny or escalate.
That is deliberately narrower than the ambition that started the project. The gap is useful: it tells us exactly which experiments still matter.
Who this is for
The book is for programmers, ML engineers, agent authors and language designers who are less interested in another prompting pattern than in the boundary between models and programs.
Python is used throughout. You do not need prior Jev experience. Familiarity with the earlier books — especially Language, Embeddings From First Principles and the agent/runtime work — helps because this book imports their ideas rather than reteaching them: representation ladders, retrieval, typed interfaces, evidence, runtime boundaries and provider separation.
Reader promise
By the end of the book you should be able to look at a decision a program needs and ask the right questions in the right order:
- Can ordinary code decide it?
- Is a trained classifier sufficient?
- Is retrieval the real bottleneck?
- What does the score actually mean?
- When must the system abstain?
- Which provider should answer this decision?
- What evidence should the decision leave behind?
- Does a proposed semantic construct add enforceable behaviour, or only prettier syntax?
You will also know what this book has not established yet, and exactly what evidence would be needed to establish it.
How to read it
If you want the argument, read Parts I, III, V and IX in order. If you are mainly interested in the proposed programming model, start with Chapter 8 for the decision contract and then read Chapters 16 to 19 and 32 to 35. If embeddings and large-corpus processing are your concern, go from Chapter 8 to Part VI. If you are evaluating whether Jev-style models are worth deploying, read Chapters 2 to 7, then Part VIII, and finish with Chapter 35 before drawing a conclusion.
The book is a research artifact as much as a manuscript. Read the negative results, the PENDING_RUN boxes and the limitations with the same weight as the successful examples. They are part of the answer.
Chapters
Stop Generating
When software needs one bounded value, what does it cost to get it by generating prose? A measured baseline before any decision model appears.
The Jev Provocation
What Jev's public contract actually says, reconstructed from the vendor's own documentation — and which of its seven properties, if any, is genuinely new.
The Jev Controversy
What the critics and defenders actually claim, checked against the one independent benchmark in full — and the frozen predictions plus the benchmark protocol that separate them.
The Smallest Decision
What deliberately boring baselines achieve on fixed-label intent and safety tasks — and the measured bar every more sophisticated provider must clear.
Zero-Shot Decisions
What runtime-defined labels gain on unseen intents, what they lose on seen ones, and how much the wording matters.
Decisions as Entailment
Whether an off-the-shelf NLI model matches a decision model on runtime-defined labels, and where the entailment formulation breaks.
The First Token
What tokenisation does to option scoring: which readouts exist at 77 options, which are impossible, and what that does to the claim that a decision model is just a way of asking a model.
The Decision Contract
A provider can answer, refuse, or violate the contract. Fake-model swaps test which distinctions the caller actually needs.
Confidence Is Not Probability
Measure what a provider's scores mean before your program branches on them.
Calibration
Spend labels on a probability map, then measure what survives a change in distribution.
Abstain
A decision may refuse to decide. You can quote what refusing costs, and you must not trust the quote off-distribution.
Unknown
Separate 'I cannot tell' from 'none of these answers' and 'this question has no answer here' before asking whether providers can separate them.
Reading the Hidden State
Ask whether a linear probe can read a decision from hidden activations more cheaply than readout, then require controls before believing it.
Tuning an Arbiter
Define Arbiter-2 as a tuned decision model, then design the fair three-way comparison that decides whether tuning earns its cost.
Transfer Across Decisions
Design the decisive test of general decision capability: diverse tasks versus relabeled volume, with H3 held at insufficient evidence until it runs.
The Decision Expression
Give the decision a calling convention: one total expression that returns a value with provenance or a typed refusal, and say what == means on it.
if decide
Branch on a decision without lying: mandatory uncertain arms, the silent failure of boolean coercion, and thresholds with statistical backing.
Semantic Match
Match on a decision with exhaustive arms, honest ties, and a measured answer to whether the wording moved the outcome.
Decision Types
Which output types are genuinely different decision semantics, and which are only representations of Choice plus post-processing?
Decisions Cannot Read Everything
What happens to a decision when the state is larger than the model can use well?
Embeddings as Candidate Generation
How do retrieval errors and decision errors combine when a decision can only read a candidate set?
where decide
Is semantic filtering a construct, or retrieval plus a decision?
for decide
Does semantic iteration deserve its own syntax, or is it sugar for filtering?
Decisions About Decisions
How does uncertainty propagate when one decision consumes another?
Semantic Pipelines
Can a real task be built from retrieval, relevance, relation, generation and support decisions?
Decision Graphs
What must be recorded so a graph of decisions can be inspected, replayed and audited?
Decisions Over Time
Do several questions over one state share work, and does that matter?
One Contract, Many Models
Where does each kind of provider win, on the same decisions?
The Decision Router
Can a router choose the cheapest competent provider for each request?
Cascades
Does uncertainty-driven escalation beat always using the strongest provider?
Learning the Router
Can a router learn which provider to use, from the outcomes it sees?
The Decision Specification
What must a declarative decision specification say, so a validator can catch mistakes the library API cannot?
The Decision Compiler
Can a compiler pick the cheapest provider that satisfies a decision contract?
Semantic Programs
What does a program with deterministic and semantic computation together look like, end to end?
Programming with Decisions
On the evidence, what is a decision model, and is decision a programming primitive?