← Pi Agents

The Context Window Is a Budget

What Pi sends to the model each turn, the arithmetic that triggers compaction, and how to tune reserveTokens and keepRecentTokens for your model.

You are halfway through a refactor. Pi has read forty files, run the test suite twice, and rewritten the same module three times because each attempt turned out to conflict with the last. Nothing is wrong. The transcript simply filled up, and the oldest parts of it stopped being sent to the model.

Most people meet this as a mood: Pi got worse at the task. It is easier to reason about as arithmetic. Every model request costs a fixed slice of a fixed window, and Pi spends that slice in named places you can read, count, and change.

Wrong: “The context window is a container. The conversation is in it, and the model reads the container.”

Correct: “The context window is a budget. Everything in it competes, and the oldest entries lose first.”

The composition of one request

Per how-pi-works.md, a model request is assembled from four parts: the system prompt, the active branch, available tools, and model settings. On top of that sit skills and context files.

    flowchart TD
  SP["system prompt<br/>base instructions and context files"] --> CTX
  AB["active branch<br/>one path through the tree"] --> CTX
  TL["tool definitions<br/>for active tools only"] --> CTX
  SK["skill descriptions<br/>always present"] --> CTX
  SI["skill instructions<br/>only when loaded on demand"] --> CTX
  CTX["contextTokens"] --> CMP{"over the line?"}
  RSV["reserveTokens<br/>16384 by default"] --> CMP
  CMP -->|no| REQ["request sent"]
  CMP -->|yes| CP["compaction"]
  CP --> REQ
  

The full skill body enters only when a task matches its description, which is what stops a directory of skills from costing you context before you have used any of them.

Two details follow directly. usage.md adds that the footer shows current context usage, and /session reports the session file, its ID, message count, token usage and cost. Both are your instruments. Read them before you guess.

The active branch is not the session file

A session is a tree. Each entry has an id and a parentId, and the current entry identifies the active branch — the path from the root to the current leaf. Only that path is sent. Everything you explored through /tree and abandoned stays in the file and stays out of the request.

Wrong: “The session is what the model sees.”

Correct: “The model sees the active branch. The session file is bigger than that, and holds every alternative you abandoned.”

This is the cheapest context saving available to you and it costs nothing. When an approach stalls, open /tree, select an earlier user message, edit it, and submit. You get a fresh branch that does not carry the failed attempt.

The trigger is arithmetic, not judgment

compaction.md states the condition exactly:

contextTokens > contextWindow - reserveTokens

reserveTokens defaults to 16384. It is the room left for the model’s own response. settings.md confirms the default and calls the setting compaction.reserveTokens, configurable in ~/.pi/agent/settings.json or <project-dir>/.pi/settings.json.

Pi checks this between turns, after tool results are appended and before the next assistant response. It also checks before a new user prompt. So the threshold is enforced against the projection it will really send, not against an estimate taken earlier.

There is one asymmetry worth knowing. A provider context-overflow error, or an early stopReason: "length", can select one compact-and-retry recovery attempt. That recovery is a repair path, not a strategy. message-types.md lists "length" as a terminal stop reason on AssistantMessage, alongside "stop", "toolUse", "error", "aborted", and "deferred".

What survives a compaction

Compaction is a summary, not a deletion. Five steps, from compaction.md:

  1. Find the cut point by walking backwards, accumulating token estimates until keepRecentTokens is reached. The default is 20000.
  2. Extract the messages from the previous kept boundary, or session start, up to the cut point.
  3. Generate a summary, passing the previous summary as iterative context when one exists.
  4. Append a CompactionEntry with the summary and firstKeptEntryId.
  5. Rebuild context as summary plus the messages from firstKeptEntryId onwards.

The original entries remain in the session tree. session-format.md spells out the persisted shape:

{"type":"compaction","id":"f6g7h8i9","parentId":"e5f6g7h8","timestamp":"2024-12-03T14:10:00.000Z","summary":"User discussed X, Y, Z...","firstKeptEntryId":"c3d4e5f6","tokensBefore":50000,"systemMessage":{"role":"system","content":"You are a coding assistant.","toolsAdded":[],"timestamp":1733235000000}}

tokensBefore is recalculated from the rebuilt, context-edited projection, so it reports the context that was actually replaced rather than the number that triggered the check.

Cut points respect tool calls

Valid cut points are user messages, assistant messages, BashExecution messages, and custom messages. Pi never cuts at tool results: a result must stay with the call that produced it, or the provider rejects the sequence.

A user-message span is a user message plus every turn up to the next user message. Normally compaction cuts at a span boundary. If a single span is larger than keepRecentTokens, Pi splits it at an assistant message and generates two summaries: one for prior context, one for the early part of the split span. The kept messages are then only the tail.

The summary has a shape

compaction.md gives the compaction format: Goal, Constraints & Preferences, Progress (Done, In Progress, Blocked), Key Decisions, Next Steps, and Critical Context. Branch summaries use the same format but stop after Next Steps.

Two details explain why this format works. Messages are first serialized to text by serializeConversation() — [User]:, [Assistant]:, [Assistant tool calls]:, [Tool result]: — so the model does not treat the text as a conversation to continue. And tool results are truncated to 2000 characters during that serialization, because read and bash output are typically the largest contributors to context size.

Tune the budget per model

Both token settings resolve independently: model override first, then the ordinary setting, then the built-in default. Keys are exact, case-sensitive provider/modelId values.

{
  "compaction": {
    "enabled": true,
    "reserveTokens": 16384,
    "keepRecentTokens": 20000,
    "modelOverrides": {
      "some-provider/big-model": { "reserveTokens": 400000 }
    }
  }
}

For a model with a 1M context window, that override triggers compaction above 600K tokens while keeping the ordinary 20000 recent tokens. enabled stays global; it is not model-specific.

Set "enabled": false to turn off automatic compaction. /compact still works — sessions.md says disabling automatic compaction does not disable the manual command.

Give the model a better cut

/compact accepts instructions. Use them when the next phase of work does not need the same detail as the last one:

/compact Preserve the failing test names and the chosen error-handling pattern. Drop the file-listing digressions.

This is the manual lever the built-in threshold cannot give you. Automatic compaction is generic; a manual compaction can be aimed.

What compaction does not do

It does not shrink the system prompt, it does not summarize tool definitions, and it does not shorten the messages after firstKeptEntryId. If your context pressure comes from a very large AGENTS.md, from fifteen enabled tools, or from a habit of cat-ing whole files through bash, compaction is treating the symptom. Fix the source.

It also does not erase anything you can recover. Navigate to a point before the compaction entry and that entry is no longer on the path, so the original messages are sent again. The same is true of a context_edit omission, which session-format.md calls branch-relative. Everything compaction removed is still in the file, one navigation away.

Read your own budget

Before blaming the model, look:

/session

Compare tokensBefore on the last CompactionEntry against your window. If a compaction lands every few turns, keepRecentTokens is too low for how much your tools emit per turn — raise it, or reduce what the tools emit.

Next

A budget only means something against something you can go back and inspect, and so far everything in this chapter has been arithmetic about content you cannot open. The next chapter opens the file itself: what is on disk, what shape it takes, and which of its lines the model ever saw.

Chapter 5: sessions and where they live.