Compose a pipeline with the weakest component at each position
Run checkClaim() on Pi 1.0.4 under scripted responses: two requests on the cheap path, one throw when a citation is invented, and a bounded agent paid only for an insufficient verdict.
Composition does not imply agency. A claim checker needs four kinds of component — a
plain function, a typed model call, a membership check and an agent — and each one is
used only where the weaker ones cannot decide. This page shows the load-bearing parts
of the canonical examples/ch40-composition/pipeline.ts, run on Pi 1.0.4 with a
scripted provider.
Run it
The two deterministic pieces need no model at all — and one of them is the check the model’s output must pass:
// 1. A plain function. No model. Deterministic, free, testable with no script.
export function retrieve(corpus: Corpus, claim: string): string[] {
const words = claim.toLowerCase().split(/\W+/).filter((w) => w.length > 3);
return Object.values(corpus).filter((text) => words.some((w) => text.toLowerCase().includes(w)));
}
// 2. A plain function again: the check the model's output must pass.
// A citation the model invented is not in the evidence, so it fails here.
export function citationsAreReal(a: EvidenceAssessment, evidence: string[]): boolean {
return a.citations.every((c) => evidence.includes(c));
}The check runs immediately after every model call and before anything downstream:
const checked = (a: EvidenceAssessment, given: string[]) => {
if (!citationsAreReal(a, given)) throw new Error("assessment cites evidence that was not provided");
};
checked(assessment, evidence);The escalation agent exists for one reason — an insufficient verdict — and is
bounded so a model that keeps searching cannot run away:
const agent = makeResearchAgent(models, model, corpus);
let turns = 0;
let overran = false;
const limit = ESCALATION_TURNS;
const unsubscribe = agent.subscribe((e) => {
if (e.type === "turn_end" && ++turns >= limit) {
overran = true;
agent.abort();
}
});
try {
await agent.prompt(`Find more evidence about: ${claim}`);
} finally {
unsubscribe();Expected outcome
Every row is observed: pipeline.test.ts (Pi 1.0.4, scripted responses) asserts
it, and the request counts are exact for that control flow.
| Scripted behaviour | Request count | Result |
|---|---|---|
| supported on the first assessment | 2 | route: "single call"; no agent was built |
| invented citation | 1 | throws “not provided”; nothing downstream ran |
insufficient verdict |
5 | agent turns, second assessment, summary |
| invented citation on the escalation path | 4 | the same throw; the summary never ran |
| truthful citation the agent found | — | passes, checked against the corpus, not the formatted text |
| a model that never stops searching | 9 | cut off at ESCALATION_TURNS, then assessed and summarised |
| cut off having found nothing | — | throws before the second assess() — nothing new to re-decide on |
The agent is the one component that can run away, so only it carries a bound —
stated in one constant, ESCALATION_TURNS.
Mechanism and limitations
Documented: everything the components rely on — one-request assess(), the
Agent loop with tools, and step()’s maxAttempts/maxTurns (pi-ai and
pi-agent-core 1.0.4). Observed: the eight tests in pipeline.test.ts, under
scripted responses. Proposed: the design — weakest component that works, checks
after the model, extract evidence rather than trust prose, and the bound value.
A membership test catches a citation that is not in the evidence; it cannot catch a conclusion that is wrong while every citation is real. The request counts are not a prediction: a real model may take a different path.
Understand this example
Copy this prompt into your AI tool. No code runs here.
Apply this example
Copy this prompt into your AI tool. No code runs here.