<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Jev From First Principles on Programmer.ie: Modern AI programming</title>
		<link>http://programmer.ie/books/jev/</link>
		<description>Recent content in Jev From First Principles on Programmer.ie: Modern AI programming</description>
		<generator>Hugo</generator>
		<language>en-US</language>
		
		
		
		
			<lastBuildDate>Thu, 08 Oct 2026 14:00:00 +0100</lastBuildDate>
		
			<atom:link href="http://programmer.ie/books/jev/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Stop Generating</title>
				<link>http://programmer.ie/books/jev/01-chapter/</link>
				<pubDate>Mon, 05 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/01-chapter/</guid>
				<description>&lt;p&gt;Every system that answers a question with a language model does the same thing: it asks&#xA;the model to write some prose and then reads one value out of it. Sometimes the prose is&#xA;a sentence, sometimes it is JSON, sometimes it is a single word followed by four&#xA;paragraphs of explanation. The value is what the application wanted. The prose is&#xA;packaging.&lt;/p&gt;&#xA;&lt;p&gt;This chapter measures the packaging. Not philosophically — with a stopwatch and a token&#xA;counter. The question is narrow on purpose: when a program needs one bounded value, what&#xA;does it cost to obtain that value by generating prose? The answer is a number, and the&#xA;number is the whole basis for everything the rest of this book tries to build.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Jev Provocation</title>
				<link>http://programmer.ie/books/jev/02-chapter/</link>
				<pubDate>Mon, 05 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/02-chapter/</guid>
				<description>&lt;p&gt;Chapter 1 ended with a question: the predicate cannot answer half the test split, so&#xA;what is the thing that handles the rest? The vendor&amp;rsquo;s answer is a product. Before this&#xA;book builds anything, it owes you a careful reading of that product&amp;rsquo;s contract — what&#xA;it promises, what it pointedly does not promise, and which parts, if any, nobody had&#xA;built before.&lt;/p&gt;&#xA;&lt;p&gt;That last clause is the whole chapter. Everything Jev does arrives wrapped in the&#xA;claim that it is a new kind of model. The claim might be true. It is more likely to be&#xA;packaging. Packaging is not nothing — a good interface changes what gets built — but&#xA;packaging and discovery are different achievements, and confusing them is how a book&#xA;like this one lies to you. So: the contract first, from the vendor&amp;rsquo;s own pages only;&#xA;then nine older mechanisms held up against it, one property at a time; then the&#xA;verdict, including the one place where the verdict went against the prediction.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Jev Controversy</title>
				<link>http://programmer.ie/books/jev/03-chapter/</link>
				<pubDate>Mon, 05 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/03-chapter/</guid>
				<description>&lt;p&gt;Chapter 2 took the contract apart and found it was six-sevenths packaging with one&#xA;candidate distinction. This chapter asks the adversarial question: granted the&#xA;contract is mostly packaging, did Jev&amp;rsquo;s &lt;em&gt;implementation&lt;/em&gt; beat the alternatives&#xA;anyway? There is exactly one independent benchmark that can answer that, and this&#xA;chapter reads all of it — including the parts that embarrass both sides.&lt;/p&gt;&#xA;&lt;p&gt;You should know the shape of the answer before the evidence: Jev won one task&#xA;outright and lost the other by three points, while a 200-million-parameter&#xA;classifier sat three-tenths of a point behind the winner at a sixth of the latency.&#xA;&amp;ldquo;Did not dominate&amp;rdquo; is the accurate reading. &amp;ldquo;Lost&amp;rdquo; is wrong. So is &amp;ldquo;won&amp;rdquo;. The rest&#xA;of this chapter is about why those three readings keep getting confused, what each&#xA;of the book&amp;rsquo;s frozen hypotheses predicts from here, and the measurement rules that&#xA;will keep later chapters from repeating the benchmark&amp;rsquo;s own mistakes — because the&#xA;benchmark makes mistakes, and finding them is part of the job.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Smallest Decision</title>
				<link>http://programmer.ie/books/jev/04-chapter/</link>
				<pubDate>Mon, 05 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/04-chapter/</guid>
				<description>&lt;p&gt;Chapters 1 through 3 argued about a product. This chapter builds the thing the&#xA;rest of the book measures against: five deliberately boring providers, two real&#xA;datasets, and the shared harness every later chapter imports. No new model, no&#xA;hosted call, no cleverness. The question is narrow on purpose: how far do boring&#xA;providers get, and with how few labels? Whatever a sophisticated provider achieves&#xA;later has to beat &lt;em&gt;this&lt;/em&gt;, by a stated margin, at a stated cost — or it has not&#xA;earned its place.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Zero-Shot Decisions</title>
				<link>http://programmer.ie/books/jev/05-chapter/</link>
				<pubDate>Tue, 06 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/05-chapter/</guid>
				<description>&lt;p&gt;Chapter 4 ended with a fixed world: 77 labels, frozen at training time, and a&#xA;bar with a number on it. This chapter breaks that world open. In any deployed&#xA;system the label set moves — new intents appear, policies get reworded, a&#xA;product renames a category — and retraining a classifier every time is the&#xA;tax the last chapter never priced. The question is what changes when the&#xA;answer set is described in language at runtime instead of fixed at training&#xA;time: what you gain, what you lose, and what the description itself costs&#xA;you. If you have read Embeddings From First Principles you know the&#xA;mechanism (label embeddings, cosine similarity); this chapter does not&#xA;re-teach it. It measures what that mechanism is worth as a decision&#xA;provider.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Decisions as Entailment</title>
				<link>http://programmer.ie/books/jev/06-chapter/</link>
				<pubDate>Wed, 07 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/06-chapter/</guid>
				<description>&lt;p&gt;Chapter 5 opened the label set: 77 banking intents, 17 of them never trained on, and a&#xA;provider that answers by embedding similarity. The question this chapter asks is older&#xA;than that experiment. Is a &amp;ldquo;decision model&amp;rdquo; partly a rediscovery of natural-language&#xA;inference? The idea is not new — Yin et al. (&lt;a href=&#34;https://arxiv.org/abs/1909.00161&#34;&gt;arXiv:1909.00161&lt;/a&gt;,&#xA;§5) turned zero-shot classification into entailment a decade ago, and Wang et al.&#xA;(&lt;a href=&#34;https://arxiv.org/abs/2104.14690&#34;&gt;arXiv:2104.14690&lt;/a&gt;, §4.3) showed that reformulating a&#xA;task as entailment and fine-tuning lightly beats methods with 500× more parameters. The&#xA;Red Hat benchmark used &lt;code&gt;BART-large-mnli&lt;/code&gt; as its zero-shot baseline&#xA;(&lt;code&gt;research/cache/redhat-benchmark.md&lt;/code&gt;, Table 1–2), so this is also our first like-for-like&#xA;contact with their ground. If a decision model is a new category, it has to beat the&#xA;cheapest strong version of the old one.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The First Token</title>
				<link>http://programmer.ie/books/jev/07-chapter/</link>
				<pubDate>Wed, 07 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/07-chapter/</guid>
				<description>&lt;p&gt;Chapter 4 fixed a world with 77 labels and put a number on it. Chapter 5 loosened&#xA;one joint: the label set could be described at run time instead of trained in.&#xA;Both chapters compared providers on &lt;em&gt;accuracy&lt;/em&gt;. This one asks a question that&#xA;comes before any provider runs.&lt;/p&gt;&#xA;&lt;p&gt;Here is the shortest version of it. A decision model that returns a probability&#xA;over a set of options must, somehow, turn &amp;ldquo;which of these?&amp;rdquo; into a number. The&#xA;most common way is to write the options into a prompt and read the model&amp;rsquo;s&#xA;distribution over whatever comes next. If the options are printed as letters,&#xA;that is a probability over 26 possible answers. If they are printed as words, it&#xA;is a probability over strings of different lengths. If they are printed as&#xA;numbers, it is a probability over digits — and digits are not the same size in&#xA;every tokenizer. The third-party account of how Jev works that this chapter set&#xA;out to rebuild (&lt;code&gt;research/primary-sources.md&lt;/code&gt;, Victor Dibia&amp;rsquo;s write-up — &lt;strong&gt;not&lt;/strong&gt;&#xA;vendor documentation; the vendor has published no internals) describes exactly&#xA;this: a prompt that ends where the answer goes, then each option scored by its&#xA;log-probability, then a softmax across options. That is the mechanism. It is&#xA;also, as we are about to see, at least four different mechanisms depending on how&#xA;the options are printed.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Decision Contract</title>
				<link>http://programmer.ie/books/jev/08-chapter/</link>
				<pubDate>Wed, 07 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/08-chapter/</guid>
				<description>&lt;p&gt;You replace the refund predicate with a classifier. The next ticket arrives, but&#xA;your question names a label the classifier has never seen. Should the application&#xA;guess, catch an exception, or pass the ticket to a human? A distribution over the&#xA;wrong labels would look reassuring while answering a different question.&lt;/p&gt;&#xA;&lt;p&gt;&lt;a href=&#34;http://programmer.ie/books/jev/01-chapter/&#34;&gt;Stop Generating&lt;/a&gt; kept that application&amp;rsquo;s action separate&#xA;from its provider. &lt;a href=&#34;http://programmer.ie/books/jev/02-chapter/&#34;&gt;The Jev Provocation&lt;/a&gt; supplied a first&#xA;request and response shape. &lt;a href=&#34;http://programmer.ie/books/jev/04-chapter/&#34;&gt;The Smallest Decision&lt;/a&gt; and&#xA;&lt;a href=&#34;http://programmer.ie/books/jev/05-chapter/&#34;&gt;Zero-Shot Decisions&lt;/a&gt; then supplied unlike providers.&#xA;&lt;a href=&#34;http://programmer.ie/books/jev/07-chapter/&#34;&gt;The First Token&lt;/a&gt; exposed a further problem: some ways of&#xA;presenting an option set cannot be read by a particular provider. You need to know&#xA;whether it answered your request at all before asking how good its answer was.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Confidence Is Not Probability</title>
				<link>http://programmer.ie/books/jev/09-chapter/</link>
				<pubDate>Wed, 07 Oct 2026 19:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/09-chapter/</guid>
				<description>&lt;p&gt;You have a program that can approve a request or send it to a person. The provider&#xA;returns a valid distribution, the selected option has a large probability, and&#xA;your next line is &lt;code&gt;if p &amp;gt; 0.9&lt;/code&gt;. What does that comparison buy you?&lt;/p&gt;&#xA;&lt;p&gt;&lt;a href=&#34;http://programmer.ie/books/jev/08-chapter/&#34;&gt;The Decision Contract&lt;/a&gt; checks that an answer covers the&#xA;requested options and that inability is explicit. It cannot check how often a&#xA;prediction is correct. You need observations for that second question. This&#xA;chapter measures the available providers on CPU. It fits no calibration method;&#xA;&lt;a href=&#34;http://programmer.ie/books/jev/10-chapter/&#34;&gt;Calibration&lt;/a&gt; owns that work.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Calibration</title>
				<link>http://programmer.ie/books/jev/10-chapter/</link>
				<pubDate>Wed, 07 Oct 2026 22:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/10-chapter/</guid>
				<description>&lt;p&gt;You have a queue of requests and a program that branches on a provider&amp;rsquo;s score.&#xA;&lt;a href=&#34;http://programmer.ie/books/jev/09-chapter/&#34;&gt;Confidence Is Not Probability&lt;/a&gt; left you a warning: a&#xA;valid distribution can be uninformative, badly scaled, or unreliable after the&#xA;input distribution changes. You can collect labels and fit a correction. How&#xA;many labels buy useful evidence, and what does your program need to remember?&lt;/p&gt;&#xA;&lt;p&gt;The specification still defines the decision and its answer set. The provider&#xA;still supplies scores. This chapter adds a fitted map and an evidence record.&#xA;The runtime checks that they belong together. Your application chooses an action;&#xA;neither a transformed score nor a provenance field chooses its consequences.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Abstain</title>
				<link>http://programmer.ie/books/jev/11-chapter/</link>
				<pubDate>Wed, 07 Oct 2026 23:45:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/11-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;Your program branches on a provider&amp;rsquo;s score. Two chapters back, &lt;a href=&#34;http://programmer.ie/books/jev/09-chapter/&#34;&gt;Confidence Is Not Probability&lt;/a&gt; showed you that the number is not a probability of being right, and &lt;a href=&#34;http://programmer.ie/books/jev/10-chapter/&#34;&gt;Calibration&lt;/a&gt; gave you maps that fix that, inside the distribution they were fitted on. One option you have not yet given the program is the third answer: &lt;em&gt;I will not act on this one.&lt;/em&gt; A decision that is never allowed to refuse must act on its worst guesses as surely as its best. Your queue of requests contains items your provider is confident and wrong about; calibration cannot remove them, it can only price them.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Unknown</title>
				<link>http://programmer.ie/books/jev/12-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/12-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;You have taught a program to refuse. &lt;a href=&#34;http://programmer.ie/books/jev/11-chapter/&#34;&gt;Abstain&lt;/a&gt; gave the caller a typed way to decline a provider&amp;rsquo;s answer and a way to price that refusal. But one refusal shape is not enough. Consider three failures in a banking assistant. It sees a vague request to change a card limit and cannot choose between two nearly identical card intents. It sees a request to order paper checks when its menu contains only electronic banking actions. And it sees a refund instruction whose order ID is blank. The first is uncertainty among answers. The second is certainty that the answers are wrong. The third is not a decision at all: the mandatory state needed to decide is absent.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Reading the Hidden State</title>
				<link>http://programmer.ie/books/jev/13-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:30:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/13-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;You have been reading decisions from model outputs: letters, digits, brackets, full strings, entailment scores and calibrated probabilities. Those are all downstream of the same expensive object: a forward pass that turns a state into hidden activations, followed by whatever projection turns those activations into an answer. This chapter asks whether you should stop earlier. If the hidden state already contains the decision in linearly accessible form, a small probe may answer more accurately and more cheaply than the output readout. If it does not, or if the probe merely memorises its training labels, then hidden-state reading is an attractive illusion.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Tuning an Arbiter</title>
				<link>http://programmer.ie/books/jev/14-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 10:30:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/14-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;Up to now the book has mostly read decisions from frozen models. &lt;a href=&#34;http://programmer.ie/books/jev/07-chapter/&#34;&gt;The First Token&lt;/a&gt; built an option-scoring provider without running its model sweep. &lt;a href=&#34;http://programmer.ie/books/jev/06-chapter/&#34;&gt;Decisions as Entailment&lt;/a&gt; measured off-the-shelf NLI and found it losing to embeddings. &lt;a href=&#34;http://programmer.ie/books/jev/13-chapter/&#34;&gt;Reading the Hidden State&lt;/a&gt; designed, but did not run, the probe comparison. This chapter asks the next question: if you are allowed to change weights, what is the cheapest change that makes a small model good at decisions?&lt;/p&gt;</description>
			</item>
			<item>
				<title>Transfer Across Decisions</title>
				<link>http://programmer.ie/books/jev/15-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 11:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/15-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;Everything so far has been about one decision at a time. &lt;a href=&#34;http://programmer.ie/books/jev/05-chapter/&#34;&gt;Zero-Shot Decisions&lt;/a&gt; showed that runtime-defined labels buy unseen-label capability at a price. &lt;a href=&#34;http://programmer.ie/books/jev/14-chapter/&#34;&gt;Tuning an Arbiter&lt;/a&gt; designed, but did not run, the comparison that decides whether tuning earns its cost. This chapter asks the question those chapters were always pointing at: when you train on some decisions, do unseen decisions get better, or did you just learn twenty separate classifications?&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Decision Expression</title>
				<link>http://programmer.ie/books/jev/16-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 12:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/16-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;You have four chapters of machinery and no way to call it. The contract validates answers. The abstain gate prices refusal. The unknown resolver names three ways of not deciding. Each is a function you could call, but nothing says in which order, what happens when one step fails, or what the caller is allowed to ignore. So here is the failure this chapter exists to prevent, in the smallest form that still runs:&lt;/p&gt;</description>
			</item>
			<item>
				<title>if decide</title>
				<link>http://programmer.ie/books/jev/17-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 13:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/17-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;&lt;a href=&#34;http://programmer.ie/books/jev/16-chapter/&#34;&gt;The Decision Expression&lt;/a&gt; gave you a value that refuses to be ignored: an &lt;code&gt;Outcome&lt;/code&gt; that is either a &lt;code&gt;Decision&lt;/code&gt; or a typed non-answer. Now you have to &lt;em&gt;use&lt;/em&gt; it, which means branching on it. And here is the failure this chapter exists to prevent, in the form every working programmer has written at least once:&lt;/p&gt;</description>
			</item>
			<item>
				<title>Semantic Match</title>
				<link>http://programmer.ie/books/jev/18-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 14:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/18-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;&lt;a href=&#34;http://programmer.ie/books/jev/17-chapter/&#34;&gt;if decide&lt;/a&gt; gave you two arms: value and uncertain. But decisions have more than two shapes. A choice among seventy-seven intents is not a boolean, and collapsing it to &amp;ldquo;answered or not&amp;rdquo; throws away the structure the caller actually branches on: &lt;em&gt;which&lt;/em&gt; answer, how close the runner-up was, and whether the question was answerable at all. This chapter builds &lt;code&gt;match decide&lt;/code&gt;, the K+1-armed generalisation, with one arm per listed option plus the mandatory uncertain arm. It asks what each arm means when scores compete on wording rather than meaning.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Decision Types</title>
				<link>http://programmer.ie/books/jev/19-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/19-chapter/</guid>
				<description>&lt;h1 id=&#34;decision-types&#34;&gt;Decision Types&lt;/h1&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;You have a decision to make. The answer could be one of several options, or a number on a scale, or a preference between two things, or a set of options. The vendor&amp;rsquo;s contract gives you three question types: Choice, Score, and Noul. But the programming language you are building needs to know which of these are genuinely different &lt;em&gt;semantics&lt;/em&gt; and which are just representations of the same underlying thing.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Decisions Cannot Read Everything</title>
				<link>http://programmer.ie/books/jev/20-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/20-chapter/</guid>
				<description>&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from earlier measured results; its own provider experiment has not been run. What did run is a token count of a real corpus and a toy model whose output is labelled as the consequence of its assumptions.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;p&gt;You are building a decision system, and the evidence it needs lives in a corpus: documents, passages, records. The decision model can read only a finite amount of text. The contract you have been using hands the provider a single &lt;code&gt;state&lt;/code&gt; string, so the easiest design is also the most tempting one: put the whole corpus in the state and let the model sort it out.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Embeddings as Candidate Generation</title>
				<link>http://programmer.ie/books/jev/21-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/21-chapter/</guid>
				<description>&lt;h1 id=&#34;embeddings-as-candidate-generation&#34;&gt;Embeddings as Candidate Generation&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter reports a CPU experiment that ran here: an open encoder&#xA;(&lt;code&gt;bge-small-en-v1.5&lt;/code&gt;) and BM25 over a small public dataset (SciFact), on the&#xA;machine described in &lt;code&gt;evidence/environment.md&lt;/code&gt;. It also shows a model-free&#xA;walkthrough. The encoder is a relevance model, not a decision model: nothing&#xA;here is evidence about Jev or about H1-H4.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 20 ended with a claim: a decision cannot read everything. If the state is larger than the model can use, the decision layer must be fed a smaller state. Something has to choose what it reads.&lt;/p&gt;</description>
			</item>
			<item>
				<title>where decide</title>
				<link>http://programmer.ie/books/jev/22-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/22-chapter/</guid>
				<description>&lt;h1 id=&#34;where-decide&#34;&gt;where decide&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter reports a CPU experiment that ran here: a cross-encoder&#xA;(&lt;code&gt;ms-marco-MiniLM-L-6&lt;/code&gt;) and an encoder (&lt;code&gt;bge-small&lt;/code&gt;) over SciFact. A&#xA;cross-encoder is a relevance scorer, not a decision model in the Jev sense.&#xA;Nothing here is evidence about Jev or about H1-H4.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 21 measured a pipeline where retrieval found the evidence and the decision layer threw a third of it away. The decision layer was a thresholded embedding cosine, and it was weak. So the question is not whether to filter, but &lt;strong&gt;which filter, and does the filter deserve syntax?&lt;/strong&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>for decide</title>
				<link>http://programmer.ie/books/jev/23-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/23-chapter/</guid>
				<description>&lt;h1 id=&#34;for-decide&#34;&gt;for decide&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter reports a CPU experiment that ran here: the loop strategies&#xA;over the Chapter 22 SciFact candidate lists and their cached cross-encoder&#xA;scores. No language model was run in this chapter. The result is a rejection&#xA;of the construct, which the chapter is allowed to make.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapters 21 and 22 showed that something must select what a decision reads, and that the filter is where the end-to-end loss lives on SciFact. This chapter asks the next question: is &lt;em&gt;iteration&lt;/em&gt; its own construct?&lt;/p&gt;</description>
			</item>
			<item>
				<title>Decisions About Decisions</title>
				<link>http://programmer.ie/books/jev/24-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/24-chapter/</guid>
				<description>&lt;h1 id=&#34;decisions-about-decisions&#34;&gt;Decisions About Decisions&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter argues from the literature and from arithmetic. It&#xA;runs &lt;strong&gt;no provider&lt;/strong&gt;. Every number below is a consequence of a constructed&#xA;contingency table or an assumed conditional table, and the assumptions are&#xA;swept and stated next to the results. None of it is evidence about any model.&#xA;The provider experiment is deferred.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 16 kept &lt;code&gt;decide&lt;/code&gt; a pure function: it takes a state, asks a provider, returns a typed outcome. It declined to say how one decision feeds another. Chapter 22 showed why that matters: a filter decides which candidates a later decision reads, and its errors become the later decision&amp;rsquo;s errors.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Semantic Pipelines</title>
				<link>http://programmer.ie/books/jev/25-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/25-chapter/</guid>
				<description>&lt;h1 id=&#34;semantic-pipelines&#34;&gt;Semantic Pipelines&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter reports a CPU run of the pipeline&amp;rsquo;s retrieval and relevance&#xA;stages over SciFact, and the full pipeline with a &lt;em&gt;declared weak&lt;/em&gt; relation&#xA;provider. The relation (entailment) and generation providers need a language&#xA;model and are deferred. The author&amp;rsquo;s private Writer evidence is off limits;&#xA;the public claims are SciFact&amp;rsquo;s.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapters 21-24 built the pieces: candidate generation, filtering, relation, composition. This chapter asks whether they fit together into a real task. The task is claim research: given a claim, retrieve evidence, decide whether the evidence supports or contradicts the claim, produce an answer, and decide whether that answer is supported.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Decision Graphs</title>
				<link>http://programmer.ie/books/jev/26-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/26-chapter/</guid>
				<description>&lt;h1 id=&#34;decision-graphs&#34;&gt;Decision Graphs&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter builds the graph machinery and runs a fake-provider&#xA;replay study on CPU. No model runs; every provider is a deterministic or&#xA;seeded fake. The numbers measure the machinery, not any model.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 25 showed a pipeline whose every error is attributable to one stage. But a list of stages is a weak container for that attribution: it cannot say &lt;em&gt;why&lt;/em&gt; a verdict depended on some evidence and not other evidence, and it cannot be &lt;em&gt;replayed&lt;/em&gt; after the fact to check whether the world has moved.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Decisions Over Time</title>
				<link>http://programmer.ie/books/jev/27-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/27-chapter/</guid>
				<description>&lt;h1 id=&#34;decisions-over-time&#34;&gt;Decisions Over Time&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;Design draft: this chapter computes &lt;strong&gt;ANALYTICAL&lt;/strong&gt; token counts under assumed&#xA;costs. No model and no inference engine runs. Every number is a consequence of&#xA;the assumed state length, question length, per-call overhead and batch-scaling&#xA;exponent, which are swept and stated next to the results. None of it is a&#xA;measurement of Prompt Cache, SGLang, or any engine.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;A decision request in the vendor&amp;rsquo;s model is one state, several questions, answered in one call. The shared state is the selling point: ask several questions over one state, and the work is shared. Chapter 26 recorded the graph around this. This chapter asks the arithmetic question the claim rests on: &lt;strong&gt;do several questions over one state share work, and does that matter?&lt;/strong&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>One Contract, Many Models</title>
				<link>http://programmer.ie/books/jev/28-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/28-chapter/</guid>
				<description>&lt;h1 id=&#34;one-contract-many-models&#34;&gt;One Contract, Many Models&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter runs no new experiment. It &lt;strong&gt;compiles&lt;/strong&gt; the winner table from&#xA;rows in &lt;code&gt;results/ch04.jsonl&lt;/code&gt;, &lt;code&gt;ch05.jsonl&lt;/code&gt; and &lt;code&gt;ch06.jsonl&lt;/code&gt; and marks every&#xA;cell without a measured &lt;code&gt;split=test&lt;/code&gt; row as &lt;code&gt;NOT_OBSERVED&lt;/code&gt;. The number of&#xA;filled cells (17) is smaller than the number of empty ones, and that is a&#xA;result.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 8 defined the contract every provider satisfies. This chapter asks the question the contract exists for: &lt;strong&gt;where does each kind of provider win, on the same decisions?&lt;/strong&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Decision Router</title>
				<link>http://programmer.ie/books/jev/29-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/29-chapter/</guid>
				<description>&lt;h1 id=&#34;the-decision-router&#34;&gt;The Decision Router&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is a &lt;strong&gt;simulation over assumed provider profiles&lt;/strong&gt;. The provider&#xA;accuracy-by-difficulty curves, costs and signal correlation are assumed and&#xA;swept; every number is a consequence of them. No real router (RouteLLM,&#xA;Hybrid LLM) is measured. The assumptions sit next to the results.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 28 ended with a table mostly full of &lt;code&gt;NOT_OBSERVED&lt;/code&gt;. A router makes that honest gap useful: it chooses, per request, the cheapest provider that is competent for that request, so quality comes from where it is measured and cost is saved elsewhere. This chapter asks: &lt;strong&gt;can a router choose the cheapest competent provider for each request?&lt;/strong&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>Cascades</title>
				<link>http://programmer.ie/books/jev/30-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/30-chapter/</guid>
				<description>&lt;h1 id=&#34;cascades&#34;&gt;Cascades&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is a &lt;strong&gt;simulation over assumed stage profiles&lt;/strong&gt;. No cascade is&#xA;measured. Every number is a consequence of assumed accuracy curves, costs,&#xA;confidence noise, shift and human capacity, all swept and stated next to the&#xA;results. In particular: the naive product of stage accuracies does &lt;strong&gt;not&lt;/strong&gt;&#xA;equal the measured end-to-end accuracy, so any &amp;ldquo;cascades save X%&amp;rdquo; headline&#xA;below is a statement about the swept assumptions, not about real providers.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Learning the Router</title>
				<link>http://programmer.ie/books/jev/31-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/31-chapter/</guid>
				<description>&lt;h1 id=&#34;learning-the-router&#34;&gt;Learning the Router&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is a &lt;strong&gt;simulation over assumed provider profiles&lt;/strong&gt;. Every&#xA;accuracy curve, cost, signal noise and audit fraction is assumed and swept;&#xA;every number is a consequence of them. No real router (RouteLLM, LinUCB) is&#xA;measured.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Chapter 29&amp;rsquo;s router was explicit rules over metadata. This chapter asks what happens when the router &lt;em&gt;learns&lt;/em&gt; the answer from the outcomes it observes. The catch is that what a router observes is decided by the router itself. A router that only sees the outcomes of the providers it selects can be systematically wrong about the providers it avoids — and worse, its mistakes can be self-reinforcing.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Decision Specification</title>
				<link>http://programmer.ie/books/jev/32-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 09:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/32-chapter/</guid>
				<description>&lt;h1 id=&#34;the-decision-specification&#34;&gt;The Decision Specification&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is model-free: a declarative &lt;code&gt;DecisionSpec&lt;/code&gt; and a validator, with&#xA;six seeded bug classes and a reported detection rate. No model runs; the&#xA;numbers are the validator&amp;rsquo;s own outputs.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;Every chapter so far asked a provider a decision through the library API. The API checks a single call in isolation. But a decision in a &lt;em&gt;program&lt;/em&gt; has a life beyond one call: it may abstain, and the abstained request must go somewhere; it may fall back to another provider; its answer classes must each have a consequence; its thresholds must be coherent with its risk target. None of that is expressible in a single library call, and none of it is checked by the call&amp;rsquo;s own validation.&lt;/p&gt;</description>
			</item>
			<item>
				<title>The Decision Compiler</title>
				<link>http://programmer.ie/books/jev/33-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 10:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/33-chapter/</guid>
				<description>&lt;h1 id=&#34;the-decision-compiler&#34;&gt;The Decision Compiler&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is a &lt;strong&gt;simulation over supplied profiles&lt;/strong&gt;. The provider curves,&#xA;costs and the cascade&amp;rsquo;s end-to-end profile are seeded from earlier rows&#xA;(Ch04&amp;rsquo;s measured LR; Ch30&amp;rsquo;s simulation rows) or are explicit ASSUMPTIONS.&#xA;The compile-time table is arithmetic on those profiles; the runtime numbers&#xA;are a SIMULATION over the same assumed curves. No new model runs.&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;A decision contract is now declarative: Chapter 32&amp;rsquo;s &lt;code&gt;DecisionSpec&lt;/code&gt; names the&#xA;inputs, answers, risk, abstention route and fallback chain. This chapter asks&#xA;the next question: &lt;strong&gt;given the spec and a set of provider profiles, can a&#xA;compiler pick the cheapest provider that satisfies the contract?&lt;/strong&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>Semantic Programs</title>
				<link>http://programmer.ie/books/jev/34-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 11:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/34-chapter/</guid>
				<description>&lt;h1 id=&#34;semantic-programs&#34;&gt;Semantic Programs&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is &lt;strong&gt;model-free&lt;/strong&gt;: the program runs over stub providers with&#xA;hand-supplied numbers (labelled ILLUSTRATIVE) and deterministic retrieval&#xA;(BM25, Chapter 21). The generation step is a stub. The ReAct-style comparison&#xA;is a &lt;strong&gt;count over a declared task script&lt;/strong&gt; — no agent is run (no LLM is&#xA;allowed by the book&amp;rsquo;s rules).&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-problem&#34;&gt;The problem&lt;/h2&gt;&#xA;&lt;p&gt;The book has spent sixteen chapters building pieces: &lt;code&gt;decide&lt;/code&gt;, &lt;code&gt;if decide&lt;/code&gt;,&#xA;&lt;code&gt;match decide&lt;/code&gt;, &lt;code&gt;where decide&lt;/code&gt;, the spec, the compiler, the graph. This chapter&#xA;asks the question the pieces were for: &lt;strong&gt;what does a program with deterministic&#xA;and semantic computation together look like, end to end?&lt;/strong&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>Programming with Decisions</title>
				<link>http://programmer.ie/books/jev/35-chapter/</link>
				<pubDate>Thu, 08 Oct 2026 12:00:00 +0100</pubDate>
				<guid>http://programmer.ie/books/jev/35-chapter/</guid>
				<description>&lt;h1 id=&#34;programming-with-decisions&#34;&gt;Programming with Decisions&lt;/h1&gt;&#xA;&lt;blockquote&gt;&#xA;&lt;p&gt;This chapter is a &lt;strong&gt;verdict over the ledger&lt;/strong&gt;, not an experiment: one&#xA;mechanical extractor (&lt;code&gt;run_ch35.py&lt;/code&gt;) reads the metadata preregistrations and&#xA;writes &lt;code&gt;results/ch35.jsonl&lt;/code&gt;; every number below is from that file or from a&#xA;cited ledger row. No model ran. The hypothesis classification follows the&#xA;author&amp;rsquo;s floor for Session E: &lt;strong&gt;no hypothesis is decided above&#xA;INSUFFICIENT_EVIDENCE without a real measured provider.&lt;/strong&gt;&lt;/p&gt;&#xA;&lt;/blockquote&gt;&#xA;&lt;h2 id=&#34;the-question&#34;&gt;The question&lt;/h2&gt;&#xA;&lt;p&gt;On the evidence, what is a decision model, and is decision a programming&#xA;primitive? The book must answer from what it actually measured, and the answer&#xA;is allowed to be narrower than the investigation set out to build.&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
