← Pi Agents

Steering, Queuing, and Changing Direction

What happens when you type while Pi is working — steering enters after the current turn, follow-ups wait for the run to finish, and Escape returns the queue.

The agent is eleven tool calls into migrating your schema. You can see it is going to be wrong in about three more steps, and stopping it now would waste everything it has already learned about the database layer.

You type. You press Enter. And the message waits — because Pi was not going to interrupt mid-tool-call, and the moment it delivers your instruction is a specific, documented point in the loop.

Two queues, not one

how-pi-works.md puts it in one sentence: “Steering messages enter after the current assistant turn. Follow-up messages enter after the agent has finished its pending work. Aborting stops the current run and returns queued messages to the editor.”

That is the whole distinction. Steering changes the current task. Follow-up adds work after it. Both are queued; they are delivered at different moments.

In the terminal, usage.md maps them to keys:

What you want Action
Adjust the current task Type a message and press Enter
Add work after the current task Type a message and press Alt+Enter
Return queued messages to the editor Press Alt+Up
Stop the current task Press Escape

“A message sent with Enter waits until the current response and its tool calls finish, then guides the next response. A follow-up sent with Alt+Enter waits until Pi finishes the current task.”

The key binding is named app.message.followUp and defaults to alt+enter, with ctrl+q on Windows and WSL — because “Windows Terminal binds Alt+Enter to fullscreen by default.” Run /hotkeys to see what your session actually has.

What a turn is

“Steering messages enter after the current assistant turn.” A turn is defined in json.md: “A turn is one assistant response plus any tool calls and tool results produced by that response.”

So the delivery point is precise: after the assistant response has finished executing all of its tool calls, and before the next LLM call. Not after one tool. After the batch the model asked for.

This is why steering feels late sometimes and immediate other times. If the model asked for three tool calls in one response, your steering message lands after all three. If it asked for one, it lands after that one.

The compaction reference confirms this sits inside the turn scheduler rather than beside it: Pi compacts during prepareNextTurn when the threshold is crossed, “then performs the existing catch-up steering poll before turn_start.” Steering is polled at a defined point in the turn pipeline.

Drawn out, that is one steering message in flight:

    sequenceDiagram
    participant U as You
    participant P as Pi
    participant M as Model
    U->>P: steer, typed while turn 1 runs
    Note over P: queued, not yet readable by the model
    P->>M: request for turn 1
    M-->>P: one response, three tool calls
    P->>P: run all three
    Note over P: turn_end
    P->>M: next request, carrying the steering message
    M-->>P: turn 2 answers the correction
    P-->>U: agent_end, then agent_settled
  

A follow-up sent at the same moment skips every arrow above it: it is “delivered only when agent has no more tool calls or steering messages.” That is the whole of “later” — it is defined by what is still outstanding, not by a clock.

How many messages arrive at once

Two settings, both defaulting to one at a time:

Setting Type Default Description
steeringMode "all" | "one-at-a-time" "one-at-a-time" How queued steering messages are delivered.
followUpMode "all" | "one-at-a-time" "one-at-a-time" How queued follow-up messages are delivered.

From rpc-commands.md:

  • "all": deliver all steering messages after the current assistant turn finishes executing its tool calls.
  • "one-at-a-time": deliver one steering message per completed assistant turn. This is the default.

The same two modes exist for follow-ups, where "all" delivers “all follow-up messages when agent finishes” and "one-at-a-time" delivers “one follow-up message per agent completion.”

One-at-a-time is the better default for a reason that is not obvious from the setting name. If you queue three corrections and all three arrive at once, the model sees them as one undifferentiated instruction. If they arrive one per turn, each one gets a turn of model attention and can be acknowledged, satisfied, or contradicted by what happens next.

From a program

Three commands, and one of them is not what it looks like.

{"type": "steer", "message": "Stop and do this instead"}

Steering is delivered “after the current assistant turn finishes executing its tool calls, before the next LLM call.” Skill commands and prompt templates are expanded. Extension commands are not allowed — the doc says use prompt instead.

The response tells you what happened:

{"type": "response", "command": "steer", "success": true, "data": {"disposition": "queued"}}

“data.disposition is "handled" if an input handler consumed this steer, or "queued" if Pi queued it (including after a handler transformed it). It does not guarantee this message remains queued.”

That last sentence matters for a UI. queued is not a promise that the message is still pending at the moment you receive the response.

Follow-ups are the same shape:

{"type": "follow_up", "message": "After you're done, also do this"}

“Delivered only when agent has no more tool calls or steering messages.” Note the ordering that implies: steering drains before follow-ups. If you queue both, the follow-up waits for the steering to land and be answered.

And prompt has a third mode:

{"type": "prompt", "message": "New instruction", "streamingBehavior": "steer"}

“If the agent is already streaming, you must specify streamingBehavior to queue the message. If the agent is streaming and no streamingBehavior is specified, the command returns an error.” "followUp" is the other accepted value.

Its disposition has three states rather than two: "handled", "queued", or "started" — "started" meaning Pi accepted it to begin a run.

Watching the queue

Every queue change is an event, and it carries the whole current queue:

{"type":"queue_update","steering":["Change direction"],"followUp":["Summarize when finished"]}

“Both fields contain the complete current queue.” Complete, not delta — so you do not need to track what you sent against what arrived. You can render this straight into a status line.

The same data is available synchronously. get_state returns:

{
  "type": "response",
  "command": "get_state",
  "success": true,
  "data": {
    "isStreaming": false,
    "steeringMode": "all",
    "followUpMode": "one-at-a-time",
    "pendingMessageCount": 0
  }
}

The SDK exposes the same split: “steer() and followUp() expose those behaviors directly and return "queued" if the input was queued (including after an extension transformed it), or "handled" if an extension consumed it.”

Aborting without losing your text

Escape stops the run. But what happens to two queued messages?

“Aborting stops the current run and returns queued messages to the editor.”

rpc-commands.md describes the client-side implementation for that behaviour, and it is not simply abort: “Remove queued steering and follow-up messages and return their text” — clear_queue. “To implement interactive Esc behavior, send clear_queue before abort, then restore the returned text in the client editor. abort continues queued messages when they remain in the session.”

    sequenceDiagram
    participant U as You
    participant P as Pi
    Note over P: turn running, two messages queued
    U->>P: clear_queue
    P-->>U: steering and followUp text
    U->>P: abort
    Note over P: run stops with nothing left pending
    U->>P: text restored into the editor
  

The ordering is the whole point. clear_queue first takes the text out; abort then stops the run with nothing pending. Abort first and, if the queue survives in the session, you recover the text some other way.

Extension commands run immediately

One asymmetry catches people writing RPC clients. From rpc-commands.md: “If the message is an extension command (e.g., /mycommand), it executes immediately even during streaming. Extension commands manage their own LLM interaction via pi.sendMessage().”

So /guard status runs while the agent is mid-run, but a steering message containing /guard is rejected. The reason is structural: an extension command is a UI action, and a steering message is model input.

When the work is actually finished

This is the trap in every client you will write.

Wrong: “agent_end means the run is over.”

Correct: agent_end closes one low-level agent run, and “Automatic retry, overflow recovery, compaction retry, steering, or follow-up work can still continue.” extensions.md names turn_end and agent_before_settle as “actionable boundaries” whose handlers can chain entries and return continue: true for one more model request. agent_settled declares no result type — it is notification-only — and its documented meaning is “Pi will not continue automatically through retries, compaction recovery, or queued messages.”

rpc.md says the same to protocol authors: “Wait for agent_settled when the client needs to know Pi will not continue automatically.”

If you print a final answer on agent_end and then a steering message arrives, you have printed something incomplete. If you declare success on agent_end and a retry follows, you have declared success twice.

Choosing what to type

Two questions. Is what you want to say a correction to what is happening now, or an item for afterwards? A correction goes with Enter — it lands inside the run and steers the next response. An item goes with Alt+Enter and waits until the run has nothing left to do.

If both are true — stop it now and have follow-up work — send the steering message, let it land, then queue the follow-up. The scheduler drains steering first.

Typing while the agent works feels like an interruption and is not one. If you need the current run to stop now, that is Escape, not a message.

That is the run loop from the outside. What it carries is the next question, because a queue that delivers the right message to the wrong point is still a lost instruction — and the shapes it delivers are not the shapes you assume.