A Stochastic Step Is a Typed Function
An agent loop decides for itself when it is finished; a typed step takes that decision away by making the only way out a call that must match a schema, then chains steps as ordinary functions.
An agent’s answer is prose, and prose is the wrong thing to build on.
You cannot pass a paragraph to the next function. You cannot branch on it without parsing it, and parsing it means guessing. The model finished because it decided it was finished, in whatever words it chose. Everything downstream inherits that guess.
Chapter 36 fixed this for a single call. This chapter fixes it for a loop. The result is the most reusable idea in the book, and it fits in one function.
The shape
A stochastic step is a function whose input is text and whose output is a value you declared:
step(models, model, system, input: string, schema) -> Promise<Static<schema>>
models and model are how the step crosses the layer boundary: they are the
pi-ai collection and the model the agent core will call. The input is text.
The typed value exists only on the way out.
The model inside it can do whatever it likes, including calling tools, changing
its mind and being wrong. What it cannot do is return anything that does not match
schema. Either the caller gets a valid value, or the caller gets an exception.
There is no third outcome in which a half-valid paragraph leaks out. That is a
statement about the shape of what comes back, and the exception carries a
kind so that a caller can tell a provider failure from a model that did not
comply.
The implementation
import { Agent, type AgentTool } from "@earendil-works/pi-agent-core";
import type { Model, Models, Static, TSchema } from "@earendil-works/pi-ai";
export interface StepOptions {
/** How many invalid submissions to tolerate before giving up. */
maxAttempts?: number;
/** How many model turns to allow in total, whatever they do. Bounds a model that never submits. */
maxTurns?: number;
extraTools?: AgentTool[];
}
// `kind` says which bound or failure ended the step, so a caller can retry a provider
// failure and not a model that answered in prose.
export type StepFailure = "provider" | "invalid-submissions" | "turn-limit" | "no-submit";
export class StepFailed extends Error {
readonly kind: StepFailure;
constructor(message: string, kind: StepFailure) {
super(message);
this.kind = kind;
}
}
// One stochastic step as a function: text in, a value that matches `schema` out.
// The model finishes only by calling `submit`. The agent core validates the
// arguments before execute() runs, so execute() only ever sees a valid value.
export async function step<S extends TSchema>(
models: Models,
model: Model<any>,
system: string,
input: string,
schema: S,
options: StepOptions = {},
): Promise<Static<S>> {
const maxAttempts = options.maxAttempts ?? 3;
const maxTurns = options.maxTurns ?? 10;
let result: Static<S> | undefined;
const submit: AgentTool = {
name: "submit",
label: "Submit",
description: "Submit the final answer. This ends the task. Call it exactly once.",
parameters: schema,
async execute(_id, params) {
// The first valid submission is the answer. A later one cannot
// overwrite it, so a model that submits twice cannot retroactively
// change the value this call returns.
if (result === undefined) result = params as Static<S>;
return { content: [{ type: "text", text: "accepted" }], details: undefined, terminate: true };
},
};
const agent = new Agent({
initialState: { systemPrompt: system, model, tools: [submit, ...(options.extraTools ?? [])] },
streamFn: models.streamSimple.bind(models),
});
// An invalid submission comes back to the model as an error result, so the
// loop re-asks by itself. Bound it, because a model can repeat a mistake.
// maxAttempts bounds invalid submissions only. A model that never calls submit
// produces no invalid submission, so the turn count is the bound that covers it.
//
// maxTurns bounds the whole run, so it must fire whether or not a submission
// has been stored. It is not conditional on `result === undefined`, and that
// is the load-bearing detail: submit returns `terminate: true`, but a batch
// only ends the turn when every completed result in it agrees (chapter 17).
// A response carrying submit *and* a non-terminating tool therefore stores a
// result and keeps looping, and a bound that switched off once a result
// existed would stop holding exactly then.
let failures = 0;
let turns = 0;
let hitTurnLimit = false;
agent.subscribe((e) => {
if (e.type === "tool_execution_end" && e.toolName === "submit" && e.isError && ++failures >= maxAttempts) {
agent.abort();
}
if (e.type === "turn_end" && ++turns >= maxTurns) {
hitTurnLimit = true;
agent.abort();
}
});
await agent.prompt(input);
// A stored submission is already schema-valid, so it is the answer. Returning
// it after enforcing the bound is deliberate: the bound is there to stop the
// loop, not to reject a value the model already committed to.
if (result !== undefined) return result;
if (failures >= maxAttempts) throw new StepFailed(`no valid submission after ${failures} attempt(s)`, "invalid-submissions");
if (hitTurnLimit) throw new StepFailed(`no submission after ${turns} turn(s)`, "turn-limit");
// A failed request is not a model decision. It ends the run with the error on the final message.
const last = agent.state.messages.findLast((m) => m.role === "assistant");
if (last && (last.stopReason === "error" || last.stopReason === "aborted")) {
throw new StepFailed(`model call failed: ${last.errorMessage ?? last.stopReason}`, "provider");
}
throw new StepFailed(failures ? `no valid submission after ${failures} attempt(s)` : "the model finished without calling submit", failures ? "invalid-submissions" : "no-submit");
}
It is short because the agent core already does most of the work. Three mechanisms are doing the job, and each one is documented.
The only way out is a tool. The model is told that submit ends the task. Its
parameters are the schema you passed in, so what the model is asked to produce is
exactly the type the function returns. The TypeScript type and the runtime schema
come from one TypeBox definition, so they cannot disagree.
Arguments are validated before execute() runs. The README says the core
parses and validates arguments before beforeToolCall and execution. So the
result = params line only ever sees a valid value. You do not write the
validation. You write the schema.
terminate: true ends the run on the submission. A tool result can carry
terminate: true as a hint that the loop should skip its automatic follow-up
request. The core honours it only when every finalized tool result in the batch
sets it, so if the model submits alongside another call, the loop continues
normally. Without it, a successful submission would be followed by one more
request in which the model says “Done!” and you pay for the sentence.
What the tests establish
Fourteen tests, run against pi-agent-core 1.0.4. Each pins one property. The last
six rows exist because an independent review found that the first version of
step() got them wrong, and the first of those six was a promised bound that the
code did not actually keep.
| Test | What it shows |
|---|---|
| Valid submission | One request, not two. terminate really did save the follow-up |
| Invalid, then valid | The loop repaired itself: the bad submission came back to the model as an error result and it tried again |
| Repeatedly invalid | Stopped at maxAttempts, and execute() never saw a value |
Prose, no submit |
StepFailed, not an empty or partial result |
| Two steps | Output of one, serialised, is the next step’s input text |
| A failed request | StepFailed with kind: "provider" and the provider’s message, not “the model finished without calling submit” |
| A model that never submits | Stopped by maxTurns, with kind: "turn-limit". maxAttempts does not bound this, because no invalid submission ever happens |
Prose, no submit |
kind: "no-submit" |
submit batched with a non-terminating tool |
maxTurns still holds. The submitted value is returned and the request count is exactly maxTurns |
| Mixed batch, then the run fails | A committed value survives a later provider failure, because it was already stored |
| Submitted twice | The first valid submission wins; a later one cannot overwrite it |
| Invalid in a mixed batch | Costs one attempt, not one per tool in the batch |
| Invalid in mixed batches, exhausted | The attempt budget still ends the run with kind: "invalid-submissions" |
Three ways to spend the same work
A typed step is not one shape. The differences are in who decides the step is over, and those differences are exactly where the bounds go wrong.
| Shape | Who ends it | Typical request count | The failure mode |
|---|---|---|---|
| One request (chapter 36) | the request ends | exactly 1, or it did not work | Returns prose, or an error value. No value at all |
| Bounded loop (this chapter) | submit, or a bound |
1 to maxTurns |
Answers in prose; submits something invalid; submits in a batch with a non-terminating tool |
| Open agent (chapter 37) | the model decides | unbounded without a limit | Keeps working. Or keeps searching. Nothing stops it but you |
flowchart LR
I["input"] --> R1["request 1"]
R1 -->|"tool calls"| X["execute all"]
X -->|"every result terminate"| E["done"]
X -->|"not all agree"| R2["request 2"]
R2 -->|"turn_end"| B{"turns >= maxTurns?"}
B -->|no| R1
B -->|"yes"| AB["abort"]
E --> V["schema-valid value or StepFailed"]
AB --> V
The R2 -->|turn_end| B edge is the one that was broken. The bound was reached by
counting turns, and then only checked whether a result existed — so a submission in a
mixed batch switched the bound off. The arrows into V are the chapter’s actual
contract: two exits, and one of them is an exception.
The second row of the table above is worth a closer look, because nothing in step()
implements retry. An invalid submission is an isError tool result. The model is shown the
error, and models usually fix the argument that was wrong. That is the loop doing
its ordinary job, pointed at your schema.
The third row exists because a model can repeat a mistake. Without a bound, a model
that keeps submitting the same invalid value would loop until a request limit or a
bill stopped it. step() counts failures from the tool_execution_end events and
calls agent.abort(). The bound is proposed, not documented. It is a pattern,
and the number is yours to choose.
maxAttempts bounds invalid submissions only. A model that calls a lookup tool
forever never submits anything invalid, so that count never moves. The first
version of this function had no other bound: against a scripted model that never
submitted, it made 31 requests under the default maxAttempts of 3. That is why
step() also counts turn_end events and stops at maxTurns. Both bounds are
the book’s pattern, not Pi’s contract.
The same review found that a failed request used to surface as “the model
finished without calling submit”, which describes a rate limit as a decision the
model made. A failed request ends the run with a final assistant message whose
stopReason is error or aborted (chapter 36’s “failure is a value”), and
step() now reads that message and throws kind: "provider". pi-ai also ships
a retry layer for exactly this failure; chapter 40 says where it belongs.
The bound that was not a bound
The turn counter originally read if (++turns >= maxTurns && result === undefined).
That guard looks harmless and it is not: it means the turn bound stops applying the
moment a submission has been stored.
Chapter 17’s batch rule is what turns that into a real defect. submit returns
terminate: true, but the core honours it only when every finalized result in
the batch agrees. So a model that calls submit and a lookup tool in one response
stores a valid answer and keeps the loop running — and from that point the turn
limit no longer fires at all.
Observed on 1.0.4 against the unfixed code: with maxTurns: 2 and a scripted model
returning submit plus lookup in every turn, step() made seven requests and
returned the submitted value. A limit described in a comment as “how many model
turns to allow in total, whatever they do” was in fact “how many turns to allow
before the first valid submission”.
The fix is to drop result === undefined from the guard, and then to decide what a
bound means once a value exists. Three choices were available:
- Ignore the stored value and throw
turn-limit. Honest about the overrun, but it discards a schema-valid answer the model already committed to. - Keep returning it, and say so. The bound’s job is to stop the loop; whether the value is trustworthy is a separate question the schema already answered.
- Track that both happened, and report a distinct kind.
This chapter takes the second, because it is the one that composes. A caller
retrying on kind: "turn-limit" is asking “did the model fail to answer?”, and a
stored submission means it did not. The ninth row of the table above pins it.
Two further points the tests pin, both about the same batch. An invalid submission
in a batch counts one attempt, not one per tool, so mixing a lookup into a
repairing turn does not burn the budget twice. And a valid submission that is
followed by a provider failure is still returned, because result was assigned
before the failure; a stored value is a commitment the loop has already made.
Chaining
Because each step has a declared output type, the value you get back is typed, and you pass it to the next step as text, serialised by you:
const triage = await step(models, model, "Triage.", report, Triage);
const action = await step(models, model, "Write the next action.", JSON.stringify(triage), Action);
There is no framework here and none is needed. The data flowing out of each step
is a plain object, and you decide how it is written into the next prompt. It is
not typed into the next step: JSON.stringify erases the type at that seam. You can log it, store it, assert on it, branch on triage.severity,
or skip the second step altogether. The model never sees anything you did not
choose to put in the next prompt, so there is no hidden conversation carried
between steps. That is the opposite of a long chat, and it is deliberate: each step
starts with an empty transcript and the input you gave it.
Each call builds its own Agent, and result is a local variable of that call.
Two steps do not share state, so steps that do not depend on each other can run
together:
const [a, b] = await Promise.all([step(/* ... */), step(/* ... */)]);
Proposed, with a caveat: nothing in the documentation promises this is safe for every provider. Rate limits, in particular, are yours to handle.
Tools inside a step
step() takes extraTools. A step that needs to look something up gets the tool
from chapter 37 beside submit:
const out = await step(models, model, system, input, EvidenceAssessment, {
extraTools: [lookupEvidence(corpus)],
});
Now the step is a small agent whose exit is typed. The model may call
lookup_evidence as many times as it needs, and the loop ends only when it calls
submit with a valid assessment. This is the point where chapter 36’s single call
and chapter 37’s open loop meet, and it is the more useful of the two as a default.
What a type cannot give you
A value that matches the schema is not a value that is true. The schema says
confidence is a number between 0 and 1. It does not say the model should be
confident. It says citations is an array of strings. It does not say the strings
exist in the evidence.
That gap is real, and it is where the next two chapters go. A step makes the shape of an answer a contract. It does not make the content of the answer one. You close the gap with ordinary code after the step, which is deterministic and free:
if (!assessment.citations.every((c) => evidence.includes(c))) throw new Error("invented citation");
Chapter 40 does that, and uses it to decide where in a pipeline the model is trusted and where it is checked.
What this is not
It is not a multi-agent system. One agent runs, finishes, and is discarded. The
word “step” is chosen over “agent” because nothing about the composition depends
on the steps being agents. A step might be one complete() call, as in chapter 36.
It might be a loop with tools. It might be a plain function that never calls a
model. Chapter 40 mixes all three.
It is also not guaranteed to succeed. step() can fail, and it fails loudly with
StepFailed. A caller that cannot tolerate failure needs a policy: retry the
step, use a different model, escalate to a person, or stop. Which of those is
right is an application decision, and the function deliberately does not make it.
What you can and cannot claim
Documented: AgentTool, argument validation before execution, terminate
semantics (including “every finalized tool result in the batch”), abort(), and
throw-to-fail (pi-agent-core README and declarations, 1.0.4).
Observed: every row of the test table, run against 1.0.4 under scripted responses. This shows that the code handles each case. It does not show how often a real model submits a valid value on the first try. Chapter 41 is how you measure that.
Proposed: the step() abstraction, the maxAttempts bound, and running
independent steps in parallel.
Next
You can now build a typed function out of a model. The next question is not about
steps but about control: where, in the loop, can you intervene? step() used one
hook, agent.subscribe(), and used it after the fact. Some decisions have to be made
before the thing they decide has happened.
Chapter 39 maps every point at which the agent core lets you interfere, and what each one can still change.