Failures, Retries, and Recovery
Three different responses to a failed request — repeat it, repair the context and retry once, or stop — and the settings and events that tell them apart.
The provider returns 529 overloaded. Pi waits two seconds and tries again. Twice more. Then it stops and tells you the model is unavailable.
That is correct — and it is the least interesting failure Pi handles. There are three responses to a failed request, they are not interchangeable, and picking the wrong one is how an agent ends up retrying a request that could never succeed.
Three responses, three meanings
Retry. The failure was transient. The request was fine; the server was not. Repeat it.
Recover. The request could never succeed as sent, but a different request would. Context overflow is the canonical case: the prompt is too long, so shrink the context and send a different prompt.
Stop. The request was wrong, or the condition is permanent. Retrying produces the same failure.
The mechanisms Pi keeps apart, and what each one does: The mechanisms Pi keeps apart, and what each one does:
flowchart TD
F["a request fails"] --> T{"transient: overloaded, rate limit, 5xx"}
T -->|yes| R["repeat it: auto_retry_start"]
R -->|succeeds| OK["run continues"]
R -->|attempts exhausted| S["stop: auto_retry_end with success false"]
T -->|no| O{"context overflow, or stopReason length"}
O -->|yes| C["compact once, then a fresh run: compaction_end with willRetry true"]
C -->|compaction call fails| Z["summarization_retry_scheduled"]
O -->|no| X{"where did it fail"}
X -->|tool execute| E["error result the model can read"]
X -->|tool_call handler| B["tool blocked as a fail-safe"]
X -->|neither| S
Three retry-shaped mechanisms with three distinct event families, and a fourth case where the answer is not a retry at all. Nothing in the system chooses between the first two for you, which is why the classification step at the end of this chapter is worth writing carefully.
Retry
rpc-commands.md defines the scope: “Enable or disable automatic retry on transient errors (overloaded, rate limit, 5xx).” Three categories, all of which mean try again in a moment.
{"type": "set_auto_retry", "enabled": true}
{"type": "abort_retry"}
The second command — “Abort an in-progress retry (cancel the delay and stop retrying)” — exists because the delay between attempts is exponential and can reach a minute. retry.maxAgentDelayMs defaults to 60000.
The agent-level settings:
| Setting | Default | Description |
|---|---|---|
retry.enabled |
true |
Enable automatic agent-level retry for transient failures. |
retry.maxRetries |
3 |
Maximum agent-level retry attempts. |
retry.baseDelayMs |
2000 |
Initial exponential-backoff delay in milliseconds. |
retry.maxAgentDelayMs |
60000 |
Maximum agent-level retry delay in milliseconds. |
And a separate provider-level layer:
| Setting | Default | Description |
|---|---|---|
retry.provider.timeoutMs |
httpIdleTimeoutMs |
Provider request timeout in milliseconds. |
retry.provider.maxRetries |
0 |
Provider-level retry attempts. |
retry.provider.maxRetryDelayMs |
60000 |
Maximum server-requested delay in milliseconds. |
Note the default of retry.provider.maxRetries is 0, and the documentation explains why with a warning rather than a rationale: “Keep retry.provider.maxRetries at 0 unless provider-level retries are required. Provider retries can delay Pi from handling quota and usage-limit errors itself.”
That is the whole design philosophy in one sentence. A retry inside the provider library is invisible to Pi, so Pi cannot decide that a quota error should stop rather than repeat. Two retry layers stacked means the outer one loses its judgement.
Timeouts sit next to retries because a hang is a failure with no error message. httpIdleTimeoutMs defaults to 300000 and “Set to 0 to disable.” websocketConnectTimeoutMs defaults to 15000. Turning a timeout off does not disable retries; it means a dead connection now looks like a slow one.
Recovery
Recovery is a different mechanism with a different trigger and a hard limit.
compaction.md: “A provider context-overflow error or an early final stopReason: "length" can select one compact-and-retry recovery attempt.”
One attempt. The context-overflow error means the request exceeded the window — sending it again, unchanged, will overflow again. The only way forward is a smaller request.
The ordering matters and is specified step by step:
persist final assistant response
→ extension/public turn_end
→ extension/public agent_end
→ append context_edit omissions for the selected attempt
→ for overflow/length: run session_before_compact and append compaction on success
→ start the retry as a fresh run
Three things to notice. The attempt that failed is still visible — turn_end and agent_end fire for it normally, and “Raw transcript history, exports, billing totals, and history-search extensions can still inspect the omitted attempt.” Recovery repairs the model context, not the record. Second, the omissions are context_edit entries, so they are branch-relative and append-only: the raw entries stay, they just stop reaching the model. Third, the retry is “a fresh run.” Not a resumption.
When recovery does not work: “If recovery compaction fails or is cancelled, Pi keeps the omission edits, appends no compaction, and schedules no internal retry.” Omissions kept, no compaction, no retry — the session ends with a repaired-but-summariless projection and the failure visible to you.
agent_before_settle “sees the repaired projection after recovery processing.” So an extension that inspects context at settle time sees what the retry will send, not what failed.
Events that tell them apart
json.md gives a distinct event family to each. This is how you tell what Pi is doing without watching the transcript.
Retry:
{"type":"auto_retry_start","attempt":1,"maxAttempts":3,"delayMs":2000,"errorMessage":"529 overloaded"}
{"type":"auto_retry_end","success":true,"attempt":2}
“On final failure, auto_retry_end has success: false and a finalError string.”
Summarization retry is a separate family, for the compaction call itself:
{"type":"summarization_retry_scheduled","attempt":1,"maxAttempts":3,"delayMs":2000,"errorMessage":"terminated"}
{"type":"summarization_retry_attempt_start","source":"compaction","reason":"threshold"}
{"type":"summarization_retry_finished"}
“For a branch summary, source is "branchSummary" and reason is absent. The reason on a compaction retry is "manual", "threshold", or "overflow".”
Compaction and recovery, where reason is "manual", "threshold" or "overflow" — and overflow is the one that tells you a recovery is in progress:
{"type":"compaction_start","reason":"threshold"}
{
"type": "compaction_end",
"reason": "overflow",
"result": {
"summary": "...",
"firstKeptEntryId": "abc123",
"tokensBefore": 150000,
"estimatedTokensAfter": 32000
},
"aborted": false,
"willRetry": true
}
“Successful overflow recovery sets willRetry to true before Pi retries the prompt.” If compaction was aborted, result is absent and aborted is true. If it failed, result is absent, aborted is false, and errorMessage describes the failure.
So: auto_retry_start means repeat. compaction_start with reason: "overflow" and a compaction_end with willRetry: true means repair. auto_retry_end with success: false means give up.
Where the classification comes from
For a custom provider, custom-provider.md is explicit about the boundary you must not cross: “Pi can compact and retry after recognized context-overflow errors. If the service uses an unknown message, normalize only that provider’s overflow response to context_length_exceeded in a guarded message_end handler.” And then, just as firmly: “Do not rewrite rate limits or transient provider failures as context overflow. Those failures use Pi’s normal retry behavior instead.”
pi.on("message_end", async (event, ctx) => {
const msg = event.message;
if (msg.role !== "assistant") return;
if (msg.stopReason !== "error") return;
if (!msg.errorMessage?.includes("model is too long")) return;
return {
...msg,
errorMessage: "context_length_exceeded",
};
});
This is the single most consequential classification in the system. Get it wrong in one direction and Pi retries an impossible request three times before failing. Get it wrong in the other and Pi silently compacts your conversation because a rate limit happened to contain the word “length”.
Failures inside extensions
The retry machinery is only about model requests. Tool and handler failures follow different rules, and extensions.md draws the line:
“Pi reports handler errors and continues where possible. A tool_call handler failure blocks the tool as a fail-safe; a tool execution failure becomes an error result for the model.”
Two different destinations. A tool_call failure blocks the tool — it never reaches the model as data. A tool execution failure becomes a result the model can read and respond to, and “Throw from execute() to produce a failed tool result” is the documented way to fail deliberately. For state you want to carry even when a result carries data: “return the result with isError: true instead of throwing: the model sees an error, and scripts still receive structuredContent.”
Shell commands are the third case. rpc-commands.md documents excludeFromContext on bash: “Set excludeFromContext to true when the command output should be stored in the session but omitted from the model context on the next prompt.” A build log nobody needs to see again is not a failure, but it is the same shape of decision — keep it out of what the model reads.
Cleanup after a failure
extensions.md: “Release resources in session_shutdown even when normal operation attempted cleanup. Keep cleanup idempotent because cancellation, reload, session replacement, and process exit can converge on the same path.”
Four different ways to end, one cleanup function. Idempotent is the requirement, not tidiness.
And “Use ctx.shutdown() to request an orderly process shutdown.” That is the difference between stopping and crashing — after a failure, a clean shutdown runs your cleanup, and kill does not.
Configuring it
A reasonable starting point, with the reasoning visible:
{
"retry": {
"enabled": true,
"maxRetries": 3,
"baseDelayMs": 2000,
"maxAgentDelayMs": 60000,
"provider": {
"timeoutMs": 300000,
"maxRetries": 0,
"maxRetryDelayMs": 60000
}
},
"httpIdleTimeoutMs": 300000,
"transport": "auto"
}
Four decisions encoded. Agent-level retry on, because the three transient categories are genuinely transient. Provider-level retry off, for the documented reason. Timeouts finite, so a dead connection becomes an error rather than a wait. transport — "auto", "sse", "websocket" or "websocket-cached", default "auto" — is worth setting explicitly for a long-running session, since a transport that reconnects well and one that does not produce very different retry experiences.
When a retry is the wrong answer
If a request fails deterministically — a bad API key, an unknown model, a malformed endpoint — three retries cost you a minute and change nothing. set_auto_retry with enabled: false, or retry.enabled: false, is the right response to a configuration error you have already diagnosed. models.md’s troubleshooting section is the same idea in prose: “Confirm that its provider has usable authentication,” and “Check its API type and compatibility settings in models.json. The upstream server must support the corresponding request fields and behavior.” None of those get better by being tried again.
The shape of it
Every failure is answered by one of four questions in order: was this transient, in which case retry with backoff. Was it deterministic but repairable, in which case compact and retry once. Was it deterministic and unfixable by Pi, in which case stop and report. Was it inside a tool or a handler, in which case the answer is not a retry but an error result the model can reason about.
That is the end of Part III. You have the run loop, the queue that redirects it, the tree that rewinds it, the messages it carries, the file that stores them, the instructions that outlive them, and the failure paths that end it. None of that tells you which one just misbehaved in your extension — which is a different question, and a mechanical one.