← Pi Agents

Authority, Isolation, and What Makes an Agent Worth Building

The closing argument — capability is not authority, a declared risk cannot gate itself, project trust is not a boundary, and how to look at a requirement and know which Pi mechanism it needs.

Everything so far has been a mechanism. Tools, extensions, events, sessions, modes, protocols. This chapter is about the one question that turns a set of mechanisms into something you would let near a production repository — and about what you are then able to claim.

Two claims carry the chapter. Capability is not authority: Pi can do almost anything the account that started it can do, and several documented mechanisms bear on whether it should do a particular thing. None of them is a sandbox, and all of them are real. And almost nothing in this book is an authority boundary except an operating-system boundary. Everything else is a procedure, a gate, or a convention — and a procedure can be designed well while remaining a procedure.

What Pi refuses to call a boundary

security.md states the position in its first paragraph:

Treat model-generated commands and code as untrusted. Pi can read, change, and execute files with the permissions of the account that started it, and it does not ask for approval before every tool call.

usage.md repeats it from the user’s side: “Pi does not ask before every tool call. Review commands and changed files, and use a sandbox for untrusted or unattended work.”

Then the sentence that should govern every design decision from here:

Watching the transcript, using project trust, and reviewing changes do not create a security boundary.

Which points at where safety actually comes from — “limiting the files, credentials, processes, and network services Pi can access and affect if a generated action is wrong or hostile.” Safety is a property of the environment, not of the transcript.

Project trust is a startup gate, not a sandbox

The mechanism people most often over-trust. security.md gives its purpose accurately — “It prevents a folder from silently loading executable extensions before you approve it” — and spends most of its section explaining what that does not buy.

What it protects

Pi requires a trust decision when the working directory supplies any of:

  • .pi/settings.json
  • .pi/mcp.json
  • .pi/extensions, .pi/skills, .pi/prompts, or .pi/themes
  • .pi/SYSTEM.md or .pi/APPEND_SYSTEM.md
  • project .agents/skills in the current directory or an ancestor directory

Granting trust allows those project settings, MCP servers, extensions, skills, prompt templates, themes, system-prompt files, and the packages they configure. “A bare .pi directory does not require project trust.”

What it does not protect — three documented gaps

The sessionDir lookup happens first. “Pi reads the project sessionDir setting while selecting or creating a session, before it resolves project trust. Declining trust prevents the remaining project settings and protected resources from loading, but it cannot undo that initial session-directory lookup.” A document that names its own exception is one you can reason about. One that claims completeness is not.

Context files load anyway. “Context files such as AGENTS.override.md, AGENTS.md, and CLAUDE.md load regardless of project trust unless you disable context loading. Treat instructions in a folder as untrusted input even when you decline project trust.” Declining trust stops the executable configuration and not the instructions — a deliberate separation, and the correct one, because an AGENTS.md cannot run code but can still steer the model.

Trust says nothing about later actions. “Project trust does not limit what tool calls can access or affect. After Pi starts, enabled tools still use the operating-system permissions of the Pi process.”

How the decision is made

A --approve or --no-approve override applies first. Then user-level and command-line extensions can handle the project_trust event, and the first to return yes or no owns it; then Pi looks for a saved decision for the current directory or a parent, closest wins; then the global defaultProjectTrust setting, whose default is "ask". Saved decisions live in ~/.pi/agent/trust.json and use canonical directory paths, and /trust saves one for future processes.

Two constraints shape who may take part. extensions.md allows only personal and explicit command-line extensions to handle project_trust, because that event runs before project extensions load. sdk.md explains why the CLI’s own builtin: entries cannot: they “load after project trust is resolved, so it cannot handle project_trust.” You cannot let the project decide whether it is trusted.

A risk a tool declares about itself cannot gate itself

The most transferable idea in this chapter, and not one Pi can give you — Pi can only give you the material.

Chapter 17 introduced tool annotations: readOnlyHint, destructiveHint, idempotentHint, openWorldHint. And it gave you Pi’s honest default, which is the part worth memorising:

Missing hints take the MCP defaults: a tool is not read-only, and may be destructive and reach an open world.

Look at where that default comes from. It does not assume a missing hint means safe. It assumes it means unknown, and resolves unknown in the dangerous direction. That is the correct resolution, and somebody had to make it deliberately.

Now take it one step further, which is not in Pi’s documentation and is yours to build:

A value chosen by the inspected party cannot gate the inspection.

If a tool, an MCP server, or a package declares its own risk level, then whoever declares it is the party you were trying to constrain. A server that finds a safety gate inconvenient can declare itself low-risk. A tool that wants to skip a confirmation can claim to be read-only. Whatever the missing-hint default is, it is defeated by a dishonest declaration rather than an absent one.

So there are two distinct protections and you need both:

Guards against Implemented as
Resolve unknown to dangerous The tool that knows nothing about itself Pi’s default for absent hints
Assign risk outside the tool’s control The tool that lies about itself Your own policy, applied to a set you control

Pi gives you the first. The second is the permission model you have to write.

Assign the risk yourself, in one place

The generalisable shape, which you can build on a tool_call handler or on the host around it:

1. CLASSIFY   look the tool up in a table you own, not one it supplied
2. EVALUATE   apply the mode, the session grants, and any hard denials
3. DECIDE     allow | confirm | deny, and at what scope
4. EXECUTE    or return a failure the model can read

Step 1 is where the authority lives, and two rules fall out of the shape that implementations get wrong.

Enforcement must not be reachable by the thing it constrains. The gate belongs outside the session the model drives, in the process that spawns the tools. A permission check the agent could influence — by writing a file the checker reads, by calling an extension that mutates the policy, or by being told its own mode — is not enforcement. It is advice.

One gate, not one per integration. The moment you have a rule for bash, a rule for the first MCP server and a different rule for the second, you have no rule. Namespace contributed tools into one space and put a single decision point in front of all of them. Chapter 17’s namespace and exposure are the levers Pi gives you for this.

Designing a gate people will not learn to wave through

A confirm dialog has exactly two states. It is astonishingly easy to build one that fires on every write and train the person using it to press the same key every time, at which point it has cost you the interruption and bought nothing. Pi gives you ctx.ui.confirm(prompt, detail) and nothing richer, so the design is yours. What it requires:

Rule What it buys
Give the decision a scope, not just a verdict. Once / for this session / never are three questions. A two-button dialog asks only the first — and “never” must not be the easy option. A grant that outlives the reason it was made also survives the project that justified it, so scope it to a session, keyed by tool.
Let the unconditional cases spend no prompt. Deny what will be denied anyway, before any card exists. Nothing to dismiss, nothing to reflexively accept.
Scope the gate to what actually varies. Reading inside the workspace is not interesting; reaching outside it is. The single largest reduction in prompts.
Carry the context in the card. Risk level, reason, arguments, path. A card without them can only be answered reflexively, because answering it properly would mean going and looking.
Queue, do not stack, and deny the queue on abort. Render only the head; match answers by request identity, not position. A late answer to an old card must never resolve a newer one, and nothing is left stranded on an answer that will never come.

That is a design, not a default. Pi documents confirm; the rest is yours to ship.

The ladder from capability to authority

Each rung is weaker than the one below it, and the ordering is the argument.

    flowchart TD
  subgraph B["BOUNDARY - limits what can be reached"]
    direction TB
    W["Rung 0: do not grant it"]
    OS["Rung 1: dedicated OS account"]
    ISO["Rung 2: container or VM"]
    W --> OS --> ISO
  end
  subgraph P["PROCEDURE - asks, observes, reviews"]
    direction TB
    GATE["Rung 3: in-process gate"]
    REV["Rung 4: review"]
    GATE --> REV
  end
  ISO --> GATE
  

Arrows point from the strongest control to the weakest. The right-hand box is exactly what security.md names — watching the transcript, using project trust, reviewing changes.

Rung 0 — do not grant the capability. pi --tools read,grep,find,ls produces an agent that can inspect but not modify; -nbt and -nt strip further. It costs one flag, and it is the strongest control available, because a capability that is not present cannot be exercised regardless of what the model decides.

Rung 1 — the operating-system boundary. Run Pi as a dedicated user. security.md tabulates what remains protected: “Anything that user cannot access.” A real boundary with a real scope, and the same table is honest that “Pi still shares the operating system and network with other users.”

Rung 2 — the isolation boundary. containerization.md tabulates four methods, and the column that matters is what is isolated:

Method What is isolated Best for
Plain Docker Pi, built-in tools, ! commands, and extensions A straightforward local container boundary
Docker Sandboxes Pi, built-in tools, ! commands, and extensions Managed isolation without exposing the real provider key
OpenShell Pi, built-in tools, ! commands, and extensions Filesystem, process, network, and credential policies
Gondolin extension Built-in tools and ! commands only A local micro-VM while retaining the host interface

That last row teaches the rest, and it is easier to see drawn than described:

    flowchart TB
  subgraph W["Whole Pi process inside the boundary"]
    HW["Host"] --> PW["Pi runtime"] --> AW["Agent"]
    PW --> EW["Extensions"]
  end
  subgraph T["Only selected tools inside"]
    HT["Host"] --> PT["Pi runtime"] --> ET["Extension tools"]
    PT --> VM["Container or VM"] --> BT["Built-in tools"]
  end
  

Gondolin keeps Pi on the host and routes read, write, edit, bash, grep, find, and ls into a Linux micro-VM. containerization.md draws the consequence: “Pi itself and other extensions remain outside the boundary, so this is a narrower form of isolation.” And: “When host Pi delegates built-in tools through Gondolin, other extension tools still run on the host unless they also delegate their work.” Isolation is not binary either. It is a question of which processes you put inside.

containerization.md gives the whole-process recipe: an image with Pi as its entry point, the working folder mounted, and a named volume for the container’s own /root/.pi/agent. Mounting the host’s ~/.pi/agent instead would expose your Pi credentials, settings, extensions, and sessions — it is on that document’s own list of what an isolated process can still reach.

Docker Sandboxes add the credential dimension: the proxy keeps the real provider key on the host and substitutes it as requests leave. containerization.md warns against the obvious mistake — “Do not run /login inside the sandbox because that writes a real credential into it.”

Rung 3 — the in-process gate. Back inside the agent: a tool_call handler that blocks produces a documented outcome, a fail-safe block the model sees as a failed result. Keep the honest label on it. This is where a declared risk meets the policy you wrote above.

Rung 4 — review. Snapshots, version control, diffs, reviewing sessions before exporting them. security.md gathers them under “Reduce impact and improve recovery”, opening with “These practices do not replace isolation.”

Confusing those two boxes is the failure mode this chapter exists to prevent — a tool_call gate that carefully confirms bash while the process it guards has unrestricted access to every credential on the machine.

Wrong model: the write was confirmed, so it was safe.

Right model: a confirmation records that a person said yes once. The capability is unchanged, and it is still the capability the whole process has.

What is not a boundary, however it is dressed

Three things that feel like controls and are not.

A timeout. Every host that interoperates with long-running code eventually puts a time budget on a hook so one slow extension cannot stall the agent. Wrong model: that budget is the boundary. Right model: it is a liveness mechanism. Trusted code can still block the thread it runs on, or do the dangerous thing quickly. A budget makes the system finish; it does not make the code safe.

A plugin sandbox around the wrong thing. Sandboxing a plugin runtime while trusted modules run inside the agent process is not isolation of the agent. The module that already holds the file and shell tools has nothing left to contain.

A reload. Re-reading a plugin’s manifest on every load is the obvious way to make development pleasant, and the obvious way to lose a control. A permission set may legitimately grow while you author a plugin; it must not be able to grow after someone approved it. The rule that protects you: a reload may narrow a permission set, never widen it. Otherwise a file write is an escalation path straight through your own gate — and the same goes for trust generally, because the permission decision is a thing you granted and anything that can edit what you granted is a way around it.

Put authority where it survives a runtime change

@earendil-works/pi-coding-agent is Pi’s package name and also the module your extension imports. That sounds like a detail. Ask a question about your own system instead:

If you replaced the agent runtime next year, which of your decisions would you have to make again?

Any decision you encoded inside the runtime has to be remade. Any decision you encoded in the layer around it does not. That is the entire argument.

For an extension author: enforcing a policy by mutating the conversation, or by steering a prompt section, makes your security decisions through a surface you do not own. The same rule enforced at an operating-system or host boundary survives the runtime changing underneath it.

For a host, the same logic prices a future runtime swap. Every authority decision made inside the runtime is re-derived and re-reviewed at that point; every decision made at the host boundary is untouched. Designing for that is not future-proofing. It is the difference between a migration and a rewrite.

Inside the runtime   ergonomics, hints, convenience, presentation
Outside the runtime  permissions, isolation, secrets, audit, refusal

If a control can be expressed in both places, put it outside. The inside copy can remain, for responsiveness, as long as the outside copy is the one that would stop a hostile action on its own.

Capability is not authority, restated as a test

Given a tool or extension:

  1. What can it reach? Files, credentials, processes, network services. Usually: everything the Pi process can reach.
  2. Who decided it may? Project trust decided whether its code loads. Nothing decided whether its actions are permitted.
  3. Who assigned its risk — you, or the tool? If the tool supplied the number, you have a claim, not a control.
  4. What constrains it? An operating-system boundary you placed, or a gate you wrote and the tool cannot influence.
  5. Would it survive a reload, or a runtime swap?

Step 4 is where every control gets its label, and the labels are the point:

Label What it means
Enforced An operating-system or virtualisation boundary stops the action without asking anyone.
Advised A gate asks a person. It can stop this call; it cannot limit what the process reaches.
Assumed A claim about the tool, the folder, or the model that nothing checks.
Outside Pi’s contract Prompt injection, and what an isolated process can still reach.

If step 4 has nothing enforced in it, you have advice.

The book in one table

The promise this book makes is that you can look at a requirement and say which Pi mechanism it needs. Here is that mapping.

The requirement says The mechanism is The contract you rely on
“It should know how we do things here” AGENTS.md / CLAUDE.md, or a skill Load order and scope; content is instruction, not code
“Do it the same way every time” A skill Loaded description; the model invokes it by name
“It needs a tool that does X” pi.registerTool() content for the model, details for rendering; throw for failure
“Do not let it do X” pi.on("tool_call") + a risk table you own Blocking is a fail-safe; a declared hint is a claim, not a control
“Don’t ask me about the obvious ones” exposure, scopes, and a mode you own Registration is not activation; annotations are unverified
“Only when I ask” ctx.hasUI / ctx.mode === "tui" JSON and print have no UI; RPC has no terminal components
“Keep it out of the model’s context” pi.appendEntry() Persisted, visible in the transcript, not in context
“It should remember between sessions” The session file, SessionManager The active branch is what reaches the model
“When this happens, do that” An event handler Order is load and registration order; some events transform, some only notify
“Start over from there” /tree, /fork, fork The branch you leave is not deleted
“Run it in CI” --print or --mode json Print signals failure by exit code; JSON reports it in the stream
“Progress, structured” --mode json, agent_settled Deltas are display-only; the end events are authoritative
“Embed it in a service” createAgentSession() SessionManager is authoritative; subscriptions rebind after replacement
“Drive it from another process” --mode rpc Response means accepted; disposition says whether a run started
“It must not touch production” An OS boundary, plus project trust for startup Trust gates loading, not actions
“It must still be safe if Pi changes” Keep permissions outside the runtime Only an OS boundary survives a runtime swap

What is still yours to decide

Nothing in this book removes these.

  • Whether a tool’s self-declared annotation is true. Pi treats a missing hint as dangerous. That is a safe default, not a guarantee.
  • Whether a given confirmation is worth interrupting a person for. A gate that prompts on everything is a gate people learn to approve blindly.
  • Which files, credentials, and network destinations the environment actually exposes. containerization.md is blunt that “an isolated process can still affect resources you expose to it.”
  • Whether the model was steered by content you did not write. security.md names prompt injection as a live concern from “files, comments, instructions, command output, and model responses,” and no configuration flag addresses it, because the content arrives through the model’s own context.
  • What to do when a load fails. One useful convention: mark a failed extension failed and step over it. A degraded agent that still runs beats a whole one that stops.

That last point is the honest limit of the whole enterprise. Everything documented in this book is a mechanism, and no mechanism can tell a malicious instruction from a helpful one, because to the model they are the same bytes in the same place.

Further reading

The mechanisms in this chapter are not invented here, and the closest worked instance is not a CLI at all. Desktop applications that embed an agent runtime publish their specs, and one of them — a published desktop host built on Pi — documents precisely this territory: a plugin trust boundary, a permission matrix with dependency rules, storage isolation, and a written plan for unbinding from the agent runtime package it started on. Read it for the mechanisms, not for the product: its choices are one team’s, several of them diverge from what Pi itself does, and where the two disagree this book follows Pi.

Closing

Thirty-five chapters ago the premise was that an agent is a model plus a loop. It is not — it is a model plus a loop, plus a working directory, plus a set of processes that can reach whatever you gave them.

You now have the vocabulary to be precise about a requirement. “It should be careful” is not a requirement. “This tool may not run without a confirmation a human sees, and the environment must not contain production credentials” is — and you know it is a tool with a risk you assigned yourself, a tool_call handler, and a container boundary, in that order of strength, with the annotation known to be unverified and the container known to be the only one of the three that is a boundary at all.

So the question this book has been circling since chapter one deserves a direct answer: what makes an agent worth building? Not capability, and not autonomy. An agent is worth building when you can state what it may reach, say who authorised that, and name every place where your own judgement stands in for a contract. That is what makes the design reviewable, and reviewability is what separates something you can hand to another person from a script with a model in it.

The discipline the book was really about is this: separate what a mechanism promises from what you have decided, and never let the second quietly borrow the authority of the first. Pi documents its own limits carefully enough that you can build on them. Everything it does not document is yours.

Build accordingly.