Example

Measure agent reliability with pass@1 and pass^k

Run eval.ts on Pi 1.0.4 against a seeded 85% simulated model: pass@1 reads 0.86, pass^5 reads 0.45, and the gap between them is the lesson.

Three activities all get called “testing”: scripted (a response you chose), structural (assert on events, not wording) and evaluation (how often the whole thing works over repeated runs). This page is the evaluation harness — the measuring parts of eval.ts, excerpted — on Pi 1.0.4. The “model” below is a seeded 85% coin, so this measures the harness’s arithmetic, not any provider; the ledger contains no real-model evidence.

Run it

The chapter’s canonical eval.ts is a self-contained file (types included); these are its load-bearing pieces. evaluate() runs every task k times, counts a thrown run as a failure, and returns the two rates:

export async function evaluate<T>(tasks: Task<T>[], k: number, options: EvalOptions = {}): Promise<EvalResult> {
	if (!Number.isInteger(k) || k < 1) throw new Error(`k must be a positive integer, got ${k}`);
	if (!tasks.length) throw new Error("evaluate needs at least one task; an empty list has no rate to report");
	const guard = options.timeout ?? defaultTimeout;
	const budget = options.timeoutMs;
	// Spread the budget over the runs, with a floor, so one run cannot consume the
	// whole evaluation's allowance.
	const perRun = budget ? Math.max(1, Math.floor(budget / (tasks.length * k))) : undefined;

	let single = 0;
	let allK = 0;
	for (const task of tasks) {
		let wins = 0;
		for (let i = 0; i < k; i++) {
			try {
				const result = perRun ? await guard(() => task.run(), perRun) : await task.run();
				if (task.passes(result as T)) wins++;
			} catch {
				// A step that throws, or times out, is a failed run. Not a crashed
				// evaluation: the point of pass^k is that a run which did not work
				// counts against the task, and a timeout is exactly that.
			}
		}
		single += wins / k;
		if (wins === k) allK++;
	}
	return { tasks: tasks.length, k, passAt1: single / tasks.length, passPowK: allK / tasks.length };
}

seeded() gives a simulated model deterministic pseudo-randomness, so the evaluation is repeatable — a real run would call the provider instead:

// Deterministic pseudo-randomness, so an evaluation of a *simulated* model is repeatable.
export function seeded(seed: number): () => number {
	let a = seed >>> 0;
	return () => {
		a = (a + 0x6d2b79f5) >>> 0;
		let t = a;
		t = Math.imul(t ^ (t >>> 15), t | 1);
		t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
		return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
	};
}

Expected outcome

testing.test.ts builds 40 tasks from one simulated model that is right 85% of the time, runs each k = 5 times with seed 42, and pins two numbers (observed):

pass@1 = 0.86
pass^5 = 0.45

pass@1 is the mean single-run rate — what a demo reports. pass^5 is the fraction of tasks that succeeded on all five runs — what a service needs. For independent runs with probability p each, all-k is p^k: 0.85⁵ ≈ 0.44, exactly what the simulation produced. Reliability compounds downward, and the average hides it. A step that throws (or times out) is a failed run, not a crash: the second evaluation test asserts the run count stays at 5 after a throw.

Mechanism and limitations

Documented: the faux provider and its response factories (pi-ai README and declarations, 1.0.4). Observed: the four tests in testing.test.ts — in particular the two pinned figures, which are properties of the simulation, not of any real model. Proposed: the three-level distinction, the rules for a meaningful evaluation (record model and date, use cases you did not tune on, include the failures that matter), and the timeoutMs budget spread across the runs.

Replacing the seeded run with a real provider is where the credential and the bill come in; neither exists here, and nothing estimates a real model’s pass^k.

Full source and test .

Understand this example

Copy this prompt into your AI tool. No code runs here.

Apply this example

Copy this prompt into your AI tool. No code runs here.