Testing Stochastic Software
Unit tests, structural tests and evaluations answer three different questions about code that contains a model; this chapter builds all three against the claim checker and shows why the average success rate hides the number that matters.
The tests in chapters 36 to 40 all passed. They tell you less than you might hope, and it is worth being exact about what.
They show that your code behaves correctly when the model does one of the things you scripted. They cannot show what the model will do. A scripted provider never refuses, never rambles, never calls the wrong tool. If you stop at green tests you have verified everything except the part that is stochastic.
Three different activities get called “testing” here, and they answer different questions. Keeping them apart is the whole chapter.
1. Scripted test Does my code handle this response correctly?
2. Structural test Did the right things happen, regardless of the words?
3. Evaluation How often does the whole thing work, over repeated runs?
The three levels, and what each can reach
flowchart TD
subgraph L1["Level 1 — scripted"]
S1["script a response<br/>assert your handling of it"] --> R1["free, deterministic, in CI"]
end
subgraph L2["Level 2 — structural"]
S2["run the real loop<br/>assert on events"] --> R2["still free under a script<br/>tolerates wording changes"]
end
subgraph L3["Level 3 — evaluation"]
S3["k independent runs<br/>over real tasks"] --> R3["costs requests, money, time<br/>not deterministic"]
end
L1 -->|"a case you thought of"| L2
L2 -->|"a rate you need"| L3
L3 -.->|"never substitute upward or downward"| X1["levels 1 and 2 never become 3"]
The solid arrows are the escalation path and they only go one way. The dotted arrow is the mistake worth naming: a green level 1 suite is routinely reported as “the agent works”, and the two sentences are about different things.
Level 1: scripted
The faux provider from chapter 36 replays responses you wrote. The test controls the model completely, so the result is exact and repeatable:
faux.setResponses([fauxAssistantMessage([fauxToolCall("submit", { severity: "high" })], { stopReason: "toolUse" })]);
assert.deepEqual(await step(models, model, "Triage.", "x", Triage), { severity: "high" });
This level is unit testing for code that happens to contain a model call. It
is the right tool for: does assess() reject a bad schema, does step() stop at
maxAttempts, does the pipeline refuse an invented citation. These are properties of
your code, and a scripted run proves them without a network, a key or a bill. It
belongs in every CI run.
Its limit is built in. The model said exactly what you told it to say. So a passing scripted test is evidence about your handling of that response, and no evidence at all that the model produces it.
A response can also be a function. fauxProvider accepts a factory that receives the
transcript the provider was sent, which is how chapter 39 checked that
transformContext changed what went out. Use that to assert on the request, not
only on the reply.
One trap: do not write assertions that only restate the script. A test that scripts
"supported" and asserts "supported" is testing the faux provider. Assert on what
your code did with the answer.
Be equally clear about what the faux provider can and cannot manufacture for you. It will happily script a refusal, a rambling answer, a wrong tool call or a malformed submission — those are the cases worth having in a suite, and they are all cases you chose. It cannot produce an autonomous surprise. Everything a scripted test exercises is behaviour you anticipated and wrote down, which makes these tests excellent at finding regressions in your handling and structurally unable to find a failure mode you did not think of. Level 3 exists for that, and levels 1 and 2 never substitute for it.
Level 2: structural
The second level asserts on what happened, not on what was said. Model wording
varies and event structure does not. A structural test runs the real loop under a
script and checks the events. Note what that does and does not demonstrate here:
under a script the wording is fixed by construction, so this test shows that the
assertions are about events, not that events stay stable when a real model’s
wording varies. That second claim needs level 3. The example below drives a small
submit agent built inline, not the claim checker, and the two further things a
structural test is good for, asserting that no tool outside an allowed set was
called and that nothing ran after a block, are proposed uses that this book has
no test for (chapter 39’s gate test measures the second one directly):
assert.ok(events.includes("tool:submit:ok"));
assert.equal(faux.state.callCount, 1, "terminate must save the follow-up request");
assert.equal(events.filter((e) => e === "turn_start").length, 1);
Assert that the right tool ran, that the loop took the expected number of turns, that no tool outside the allowed set was called, that nothing ran after a block. Those statements stay true if you swap the model, and they are the ones to carry into level 3 when you point the same checks at a real provider: the words will differ on every run, and the structure should not.
Request counts are a particularly good structural assertion, because they are a cost and a behaviour at once. “The cheap path makes exactly two requests” is a regression test for a pipeline’s economics.
Level 3: evaluation
Now the question changes. You are no longer asking whether code is correct. You are asking how often a system with a random component produces an acceptable result. That is measured, not asserted, and a single run measures nothing.
// Level 3 of chapter 41: evaluation. A scripted provider cannot tell you how
// often a real model gets it right. Repeated runs can, and the number to read
// is not the average.
export interface Task<T> {
id: string;
run: () => Promise<T>;
passes: (result: T) => boolean;
}
export interface EvalResult {
tasks: number;
k: number;
/** Mean single-run success rate. What a demo reports. */
passAt1: number;
/** Fraction of tasks that succeeded on ALL k independent runs. What a service needs. */
passPowK: number;
}
export interface EvalOptions {
/** Wall-clock budget for the whole evaluation. Per-run, not per-task, so a
* task that hangs cannot hang the evaluation past its share of the budget. */
timeoutMs?: number;
/** Injected for tests; defaults to the platform timer. */
timeout?: (fn: () => Promise<unknown>, ms: number) => Promise<unknown>;
}
const defaultTimeout = (fn: () => Promise<unknown>, ms: number) =>
new Promise((resolve, reject) => {
const t = setTimeout(() => reject(new Error(`timed out after ${ms}ms`)), ms);
fn().then(resolve, reject).finally(() => clearTimeout(t));
});
// An empty task list divides by zero and reports NaN for both rates, which reads
// as a result. A non-positive k makes pass^k vacuously 0 and pass@1 meaningless.
// Both are caller mistakes, and the numbers they produce look like measurements.
export async function evaluate<T>(tasks: Task<T>[], k: number, options: EvalOptions = {}): Promise<EvalResult> {
if (!Number.isInteger(k) || k < 1) throw new Error(`k must be a positive integer, got ${k}`);
if (!tasks.length) throw new Error("evaluate needs at least one task; an empty list has no rate to report");
const guard = options.timeout ?? defaultTimeout;
const budget = options.timeoutMs;
// Spread the budget over the runs, with a floor, so one run cannot consume the
// whole evaluation's allowance.
const perRun = budget ? Math.max(1, Math.floor(budget / (tasks.length * k))) : undefined;
let single = 0;
let allK = 0;
for (const task of tasks) {
let wins = 0;
for (let i = 0; i < k; i++) {
try {
const result = perRun ? await guard(() => task.run(), perRun) : await task.run();
if (task.passes(result as T)) wins++;
} catch {
// A step that throws, or times out, is a failed run. Not a crashed
// evaluation: the point of pass^k is that a run which did not work
// counts against the task, and a timeout is exactly that.
}
}
single += wins / k;
if (wins === k) allK++;
}
return { tasks: tasks.length, k, passAt1: single / tasks.length, passPowK: allK / tasks.length };
}
// Deterministic pseudo-randomness, so an evaluation of a *simulated* model is repeatable.
export function seeded(seed: number): () => number {
let a = seed >>> 0;
return () => {
a = (a + 0x6d2b79f5) >>> 0;
let t = a;
t = Math.imul(t ^ (t >>> 15), t | 1);
t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
Two numbers come out, and the gap between them is the lesson.
passAt1 is the mean success rate of a single run. It is what a demo reports:
“it works about 86% of the time”.
passPowK, written pass^k, is the fraction of tasks that succeeded on every one
of k independent runs. It is what a service needs: for a given input, does it work
when I call it, and when I call it again, and again?
The test simulates a model that is right 85% of the time on every task, across 40 tasks and 5 runs each, with a seeded generator so the simulation is repeatable:
pass@1 = 0.86
pass^5 = 0.45
The first number reads as healthy. The second says that on more than half of the tasks, at least one of five identical runs failed. They are not in conflict. If each run is independently right with probability 0.85, five in a row is about 0.85⁵, near 0.44, which is what the simulation produced. The average conceals that reliability compounds downward.
Be clear about what this does and does not demonstrate. The “model” is a scripted
function drawing from a seeded random generator. The test demonstrates the
arithmetic and the harness. It tells you nothing about any real model’s rate. A
real evaluation replaces the simulated run with a call to a real provider and has
no seed.
Running it against a real model
The change is in how run is built, and nothing else in the harness moves:
import { builtinModels } from "@earendil-works/pi-ai/providers/all";
const models = builtinModels();
const model = models.getModel("anthropic", "claude-sonnet-5")!;
const tasks = cases.map((c) => ({
id: c.id,
run: () => assess(models, model, c.claim, c.evidence),
passes: (r) => r.verdict === c.expected,
}));
const result = await evaluate(tasks, 5);
Documented-API, not run by this book’s suite, since it needs a credential and costs money. Four rules make the number meaningful:
- Record what you measured. The model identifier and version, the date,
k, the number of tasks, and the pinned Pi release. A pass^k with no model attached is a number with no meaning. - Use cases you did not tune on. If you edited the prompt until your twenty examples passed, you have measured your prompt’s fit to those twenty.
- Include the failures that matter. The invented citation, the contradictory evidence, the claim with no relevant evidence at all. A set made only of easy cases measures nothing you will be hurt by.
- Count a thrown step as a failed run.
evaluate()does, and a test now exercises that branch: a run that throws is counted as a failure and the evaluation carries on. (The first version’s simulated model never threw, so this branch was untested, which an independent review caught.) A step that raisesStepFailedis a result, not a crash, andStepFailed.kindlets you count provider failures separately from a model that did not comply. - Validate the inputs, because the failure is silent.
evaluate()now rejects an empty task list and akthat is not a positive integer. Both would otherwise produce numbers: an empty list divides by zero and reportsNaNfor both rates, andk = 0makes pass^k vacuously zero while pass@1 is meaningless.NaNin a results table is easy to skim past, and a vacuous zero looks like a measurement. Two more tests cover these. - Give the evaluation a deadline. A pass^k run is
tasks × krequests against a real model, and the failure mode of an unbounded one is a run that never finishes and reports nothing.evaluate()takestimeoutMs, spreads it across the runs, and counts a timeout as a failed run — which is what it is. This is the book’s pattern, not Pi’s.
The three levels together
| Scripted | Structural | Evaluation | |
|---|---|---|---|
| Question | Does my code handle this response? | Did the right things happen? | How often does it work? |
| Model | faux, fully scripted | faux or real | real |
| Deterministic | yes | yes under a script | no |
| Cost | none | none under a script | requests, money, time |
| Runs | every commit | every commit | on a schedule, and when the model or prompt changes |
| A failure means | a bug in your code | a bug, or a behaviour change | the system is less reliable, or the model changed |
| A pass means | your code handles that case | the structure held | a measured rate, with error bars you must supply |
The common failure is substitution: treating level 1 as level 3. A suite of scripted tests that all pass says “the plumbing is correct”, and teams hear “the agent works”. The two sentences are about different things, and only the second is about the model.
What the numbers do not say
Forty tasks by five runs is 200 requests and gives a pass^5 with wide error. In this chapter’s simulation the “forty tasks” are one task repeated forty times, so the 200 requests are 200 draws of one 85% coin, and the quoted figures (0.86 and 0.45, pinned exactly in the test) describe that coin, not forty different problems. With genuinely distinct tasks the spread would be different, usually wider. A difference of a few points between two prompts is probably noise at that size. If a decision rides on the difference, run more tasks, and do not read a single evaluation as a precise measurement.
A pass also means “the answer matched what you labelled correct”. That is only as
good as the label. Where several answers are acceptable, passes has to say so.
And a result describes a model at a time. Providers change models behind a stable name. A pass^k from last month is a statement about last month, which is why the record in the first rule above is not bookkeeping.
The external idea
pass^k as a way to measure agent reliability is not this book’s invention. The
reference is in Appendix A, which also flags that the figures quoted from that work
were not independently verified here. The definition used above stands on its own: it
is the code in eval.ts, and the arithmetic is checkable by hand.
What you can and cannot claim
Documented: the faux provider and its response factories (pi-ai README and
declarations, 1.0.4).
Observed: the four tests in examples/ch41-testing/. In particular that the
simulated 85% model yields pass@1 of 0.86 and pass^5 of 0.45 over 40 tasks, with
seed 42, on 1.0.4. That figure is a property of the simulation. It is not a
measurement of any real model.
Proposed: the three-level distinction, the rules for a meaningful evaluation, and the schedule in the table. These are practice, not Pi’s contract.
Next
You have a claim checker that is typed, composed, guarded and measured. What you have not asked is what it is allowed to affect. It reads a corpus. If it could also write one, send a message or run a command, none of the five chapters before this would tell you whether it should.
The last chapter is about the difference between what an agent can do and what it may.