Measure agent reliability with pass@1 and pass^k
Run eval.ts on Pi 1.0.4 against a seeded 85% simulated model: pass@1 reads 0.86, pass^5 reads 0.45, and the gap between them is the lesson.
Three activities all get called “testing”: scripted (a response you chose),
structural (assert on events, not wording) and evaluation (how often the whole thing
works over repeated runs). This page is the evaluation harness — the measuring parts
of eval.ts, excerpted — on Pi 1.0.4. The “model” below is a seeded 85% coin,
so this measures the harness’s arithmetic, not any provider; the ledger contains
no real-model evidence.
Run it
The chapter’s canonical eval.ts is a self-contained file (types included); these
are its load-bearing pieces. evaluate() runs every task k times, counts a thrown
run as a failure, and returns the two rates:
export async function evaluate<T>(tasks: Task<T>[], k: number, options: EvalOptions = {}): Promise<EvalResult> {
if (!Number.isInteger(k) || k < 1) throw new Error(`k must be a positive integer, got ${k}`);
if (!tasks.length) throw new Error("evaluate needs at least one task; an empty list has no rate to report");
const guard = options.timeout ?? defaultTimeout;
const budget = options.timeoutMs;
// Spread the budget over the runs, with a floor, so one run cannot consume the
// whole evaluation's allowance.
const perRun = budget ? Math.max(1, Math.floor(budget / (tasks.length * k))) : undefined;
let single = 0;
let allK = 0;
for (const task of tasks) {
let wins = 0;
for (let i = 0; i < k; i++) {
try {
const result = perRun ? await guard(() => task.run(), perRun) : await task.run();
if (task.passes(result as T)) wins++;
} catch {
// A step that throws, or times out, is a failed run. Not a crashed
// evaluation: the point of pass^k is that a run which did not work
// counts against the task, and a timeout is exactly that.
}
}
single += wins / k;
if (wins === k) allK++;
}
return { tasks: tasks.length, k, passAt1: single / tasks.length, passPowK: allK / tasks.length };
}seeded() gives a simulated model deterministic pseudo-randomness, so the
evaluation is repeatable — a real run would call the provider instead:
// Deterministic pseudo-randomness, so an evaluation of a *simulated* model is repeatable.
export function seeded(seed: number): () => number {
let a = seed >>> 0;
return () => {
a = (a + 0x6d2b79f5) >>> 0;
let t = a;
t = Math.imul(t ^ (t >>> 15), t | 1);
t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}Expected outcome
testing.test.ts builds 40 tasks from one simulated model that is right 85% of the
time, runs each k = 5 times with seed 42, and pins two numbers (observed):
pass@1 = 0.86
pass^5 = 0.45pass@1 is the mean single-run rate — what a demo reports. pass^5 is the fraction
of tasks that succeeded on all five runs — what a service needs. For independent
runs with probability p each, all-k is p^k: 0.85⁵ ≈ 0.44, exactly what the
simulation produced. Reliability compounds downward, and the average hides it. A
step that throws (or times out) is a failed run, not a crash: the second evaluation
test asserts the run count stays at 5 after a throw.
Mechanism and limitations
Documented: the faux provider and its response factories (pi-ai README and
declarations, 1.0.4). Observed: the four tests in testing.test.ts — in
particular the two pinned figures, which are properties of the simulation, not of
any real model. Proposed: the three-level distinction, the rules for a
meaningful evaluation (record model and date, use cases you did not tune on, include
the failures that matter), and the timeoutMs budget spread across the runs.
Replacing the seeded run with a real provider is where the credential and the bill
come in; neither exists here, and nothing estimates a real model’s pass^k.
Understand this example
Copy this prompt into your AI tool. No code runs here.
Apply this example
Copy this prompt into your AI tool. No code runs here.