Authority, Isolation, and What Makes an Agent Worth Building
The closing argument: capability is not authority, a boundary must act in time and be out of reach of what it constrains, a value chosen by the inspected party cannot gate the inspection, and where each requirement belongs on the four-rung ladder from instruction to host containment.
Everything so far has been a mechanism. Tools, extensions, events, sessions, modes, protocols. This chapter is about the one question that turns a set of mechanisms into something you would let near a production repository, and about what you are then able to claim.
Two claims carry the chapter. Capability is not authority: Pi can do almost anything the account that started it can do, and several documented mechanisms bear on whether it should do a particular thing. None of them is a sandbox, and all of them are real. And almost nothing in this book is an authority boundary except an operating-system boundary. Everything else is a procedure, a gate or a convention, and a procedure can be designed well while remaining a procedure.
The chapter earns the second claim with two tests that a mechanism must pass to be called a boundary, and with one principle about who is allowed to supply the facts a gate relies on.
Two rules, and why they point in opposite directions
Chapter 11 gave a rule for behaviour and a rule for placement:
Use the least powerful mechanism that solves the requirement. Put the decision in the narrowest layer that owns it, and only in a layer that can act before the decision becomes irreversible.
This chapter adds the rule for authority:
Enforce a prohibition at the lowest boundary that can still prevent the consequence.
They look as if they disagree. The first says climb the ladder only as far as you must, starting with a text file. The second says go as low as you can, to the operating system. They do not disagree, because they answer different questions. The behaviour rule asks what is the cheapest way to get the agent to do this? The authority rule asks what is the lowest place at which this can be made to hold, if the agent, the model and the code above it are all wrong at once?
Applying the wrong one fails in a recognisable way:
| Misapplied | What you end up with |
|---|---|
| The behaviour rule to a prohibition | “Never deploy to production” written in AGENTS.md. The cheapest mechanism, and one that cannot stop anything |
| The authority rule to a convention | A container to enforce a style guide. Strong, expensive and aimed at nothing |
Most real requirements have both halves, and they are met at different rungs. “Use our test command, and never push to the release branch” is a convention and a prohibition. The convention belongs in text. The prohibition belongs low. Chapter 11’s worked case 4 is exactly this shape, and the rest of this chapter says where each half goes and what each rung can honestly claim.
What Part V built, and what it did not
Chapters 36 to 41 built a claim checker: a model call that returns a typed value, an agent loop with a lookup tool, a typed step that cannot finish in prose, a map of every hook on the agent core, a pipeline that spends the expensive part only when the cheap part cannot decide, and a way to measure how often it works.
Every one of those is a capability. None of them is authority, and the claim checker is safe for a reason you can name: its only tool, lookup_evidence, reads a corpus you passed in and does nothing else. That is the strongest control at rung 2 of the ladder below, “do not grant the capability”, and it was the default because the core ships no tools. Chapter 39’s last row said the rest: no hook in the agent core can undo an effect, and the one hook that can refuse an effect, beforeToolCall, can only do so before it happens and only on the table you wrote.
Add a write_file tool to that agent and nothing in the five chapters before this one tells you whether it should be allowed. That is the question this chapter answers, for the coding agent and for anything you build on the core.
What Pi refuses to call a boundary
security.md states the position in its first paragraph:
Treat model-generated commands and code as untrusted. Pi can read, change, and execute files with the permissions of the account that started it, and it does not ask for approval before every tool call.
usage.md repeats it from the user’s side: “Pi does not ask before every tool call. Review commands and changed files, and use a sandbox for untrusted or unattended work.”
Then the sentence that should govern every design decision from here:
Watching the transcript, using project trust, and reviewing changes do not create a security boundary.
It points at where safety actually comes from: “limiting the files, credentials, processes, and network services Pi can access and affect if a generated action is wrong or hostile.” Safety is a property of the environment, not of the transcript.
Project trust is a startup gate, not a sandbox
The mechanism people most often over-trust. security.md gives its purpose accurately, “It prevents a folder from silently loading executable extensions before you approve it”, and spends most of its section on what that does not buy. Chapter 8 has the full account, and ran it. Three of its conclusions carry this chapter’s argument.
- Context files load anyway. Declining trust stops the executable configuration and not the instructions, so a folder can still steer the model. Chapter 8 observed it: with the project untrusted, its extension and prompt template do not load and its
AGENTS.mdstill does. A deliberate separation, and the correct one, because anAGENTS.mdcannot run code but can still steer. It is also the clearest example of why instruction is not authority: the text you did not write arrives through the same channel as the text you did. - Trust says nothing about later actions. “Project trust does not limit what tool calls can access or affect. After Pi starts, enabled tools still use the operating-system permissions of the Pi process.”
- A startup gate has an exception. The project
sessionDiris read before trust is resolved.
And you cannot let the project decide whether it is trusted. extensions.md allows only personal and explicit command-line extensions to handle project_trust, because that event runs before project extensions load.
The ladder: four rungs
Pi’s mechanisms sort into four rungs by what they can still change when the consequence is on its way. The sort is this book’s proposal. Each fact placed on it is documented or observed, and the placement is argued.
4 HOST CONTAINMENT the operating system, a container or VM, the credentials
and network the process is given at all
3 RUNTIME INTERCEPTION code that runs before an effect and can refuse it:
beforeToolCall, a tool_call handler, a gate in the host
2 AGENT POLICY what the agent is given: which tools exist, their schemas,
the mode, any risk the tool declares about itself
1 INSTRUCTION text the model reads: AGENTS.md, the system prompt,
a skill, a prompt template
Rung 4 is lowest in the sense the authority rule uses: closest to the consequence and hardest to reach from above. The numbers run the other way from “power”. An instruction is the least powerful mechanism and the cheapest, which is why the behaviour rule starts there.
Rung 1: instruction
An instruction changes what the model is likely to propose. It does not change what can execute. Two things in this book show the gap without a model: chapter 8’s observation above, that an untrusted folder’s AGENTS.md still loads, and chapter 39’s, that the only thing standing between a proposed call and its effect is code. A prompt that says “never call forbidden_action” is not in the path of execute().
Not measured here with a real model. This book has no real-model evidence yet, so whether, and how often, a model ignores a stated prohibition under pressure is not something it can tell you. The structural point does not depend on the answer: even if a model never proposed the call, the transition would still be executable.
Use instruction for what it is good at, which is steering, convention and explanation. Do not use it for a prohibition that has to hold.
Rung 2: agent policy
Agent policy is what you decide before the run: which tools exist, with which schemas. -nbt and -nt strip further, and --tools names an exact set (cli.md).
Say precisely what such a flag buys, because “which tools exist” is two different claims sitting next to each other and only one of them is about the model.
What it decides: what the model may call. --tools read,grep,find,ls declares four tools and withholds bash, edit and write from the model. A capability the model is not given cannot be called, whatever it decides, and that is why this is the strongest control within the agent. It is a real reduction and worth having.
What it does not decide: what the process can reach. Pi is a program running as you. --tools configures the model’s vocabulary; it does not change the operating-system permissions of the Pi process, and it does not constrain code that is not asking to be given a tool. Anything already loaded — an extension, a skill’s bundled script that a model has been told to run, your own host code — has the same access it had before the flag. Nothing on this rung is containment. Containment is rung 4, and the distinction is the subject of this chapter.
Two consequences follow, and both are easy to get wrong:
- The MCP gap is part of the first claim, not a caveat on it. Since 1.0.4,
--toolskeeps MCP tools unless an entry starts withmcp__. So on a machine with MCP servers configured,pi --tools read,grep,find,lsalone does not produce a four-tool agent — the servers’ tools remain registered and reachable at their exposure.--no-mcpis what makes the four-tool claim true, and it is part of the command rather than a footnote to it. Naming the servers you want and dropping the rest (--tools read,bash,'mcp__radius__*') is the selective version. - A narrower model vocabulary does not narrow a shell.
bashcan run any command the account can run, so a list omittingwritewhile includingbashhas not removed writing. The question at this rung is never “which tools did I leave out” but “what can the tools I left in reach, taken together”. Chapter 37 shows the same thing from below: the core ships no tools, so the claim checker’s only reach is the one tool you handed it.
The limit is closure, and it has a name worth using: what you have bounded is the agent’s vocabulary, not its reach. If any granted tool can do what a withheld one could, the withheld one was never withheld. Only rung 4 bounds reach.
Rung 3: runtime interception
This is the rung with the most evidence in the book, because chapter 39 measured it. A beforeToolCall that blocks a call means execute() runs zero times and the side effect is absent. A afterToolCall that rewrites the result is too late: execute() has run once. In the coding agent the same seam is the tool_call event, and chapter 16 ran a guard on it: with no UI to ask, it refuses a force push and the command never runs; a handler that throws blocks the tool as a fail-safe.
What it is, then: a gate that acts before the effect and can say no. What it cannot be is more than the call it sees. Three observations from the book mark the edges.
- It judges a representation. Chapter 11’s guard first matched
git pushwith a pattern that looked right and never fired on a plaingit push, because the handler tests the serialised input,{"command":"git push origin main"}, wheregitfollows a quote and not whitespace. It only fired for compound commands. A gate over a command string is a procedure: it blocks the spellings its author anticipated. - It judges a name. Chapter 21 observed that an MCP tool is named
mcp__<server>__<tool>, and a name-keyed gate that forgets the prefix lets it through. - It does not see past the call. A permitted
bashcall is permitted. What the process then does with the permissions it has is not the gate’s business, which is why a gate that carefully confirmsbashis no substitute for limiting what the process can reach.
Where the gate runs also matters, and this is proposed, not documented. A gate in the same process as the agent runs with the same permissions as the tools it guards. If any tool can write to something the gate reads, or to the settings and extensions loaded at the next start, the gate is advice. The stronger placement is outside the session the model drives, in the process that spawns the tools. Enforcement must not be reachable by the thing it constrains.
Rung 4: host containment
security.md tabulates how Pi can run, and the column that matters is what remains protected:
| How Pi runs | What remains protected |
|---|---|
| Directly, with the permissions of its operating-system user | “Anything that user cannot access.” A dedicated user can narrow that, but Pi “still shares the operating system and network with other users” |
| Entirely inside a container, virtual machine or sandbox | Host files and processes you do not expose; “credentials and network services remain accessible if you make them available inside it” |
| Outside the isolated environment, with only its built-in tools inside | Host resources are protected from actions performed through those tools. “Pi itself and other extensions remain outside the boundary, so this is a narrower form of isolation” |
containerization.md adds the column that makes the last row concrete, what is isolated:
| Method | What is isolated | Credential handling |
|---|---|---|
| Plain Docker | Pi, built-in tools, ! commands, and extensions |
Credentials passed into the container |
| Docker Sandboxes | Pi, built-in tools, ! commands, and extensions |
Provider credentials stay on the host and are substituted by the proxy |
| OpenShell | Pi, built-in tools, ! commands, and extensions |
Policy-controlled credentials and inference routing |
| Gondolin extension | Built-in tools and ! commands only |
Stored Pi credentials stay on the host, but commands inherit host environment variables |
That last row teaches the rest, and it is easier to see drawn than described:
flowchart TB
subgraph W["Whole Pi process inside the boundary"]
HW["Host"] --> PW["Pi runtime"] --> AW["Agent"]
PW --> EW["Extensions"]
end
subgraph T["Only selected tools inside"]
HT["Host"] --> PT["Pi runtime"] --> ET["Extension tools"]
PT --> VM["Container or VM"] --> BT["Built-in tools"]
end
Gondolin keeps Pi on the host and routes the built-in tools into a Linux micro-VM, so “other extension tools still run on the host unless they also delegate their work.” Isolation is not binary. It is a question of which processes you put inside, and of what you hand them: the same document is blunt that “an isolated process can still affect resources you expose to it.” It gives the whole-process recipe, an image with Pi as its entry point, the working folder mounted and a named volume for the container’s own /root/.pi/agent. Mounting the host’s ~/.pi/agent instead would expose your credentials, settings, extensions and sessions. Docker Sandboxes add the credential dimension, and containerization.md warns against the obvious mistake: “Do not run /login inside the sandbox because that writes a real credential into it.”
Documented, not run. This book did not run a container or a micro-VM. Its claims about this rung are the documentation’s, and the one thing it adds is the argument for why this rung is the boundary: it is the only one that limits what the consequence can be whatever ran, and nothing above it can widen it.
The cost is coarseness. A container cannot tell a good git push from a bad one. That is why it is the floor and not the whole design.
A boundary has to pass two tests
Put the four rungs next to the two things a boundary must do.
| Rung | Acts before the consequence? | Out of reach of what it constrains? | So it is |
|---|---|---|---|
| 1 Instruction | No. It changes what is proposed, not what can run | No. The model can ignore it, and content the model reads can contradict it | influence |
| 2 Agent policy | Yes, for a tool never granted | Yes, if the host builds the list, but it bounds vocabulary, not reach | a boundary on what can be called |
| 3 Runtime interception | Yes. Measured in chapter 39 | Depends on where it runs. In the agent’s process, only partly | a boundary for the calls it sees, when it sits out of reach; a procedure otherwise |
| 4 Host containment | Yes: it limits what the consequence can be | Yes: the process inside cannot widen it | an authority boundary |
The first test is chapter 1’s and chapter 39’s: can it act before the consequence is irreversible? The second is this chapter’s: can the thing it constrains reach it, or choose the facts it relies on? Only rung 4 passes both without conditions, which is the second claim of the chapter in a table. The table is proposed, built from the documented and observed facts above.
A value chosen by the inspected party cannot gate the inspection
The most transferable idea in this chapter, and not one Pi can give you. Pi can only give you the material.
A value chosen by the inspected party cannot gate the inspection.
If a tool, an MCP server, a package or the model declares a property of itself, and a gate relies on that declaration to decide whether to inspect, then the party you meant to constrain has decided how closely it is watched. It does not need to be malicious. It only has to be wrong, or to find the gate inconvenient. Three instances, and the book has run two of them.
A tool’s declared risk. Chapter 17 introduced tool annotations (readOnlyHint, destructiveHint, openWorldHint) and Pi’s default for a missing hint: “Missing hints take the MCP defaults: a tool is not read-only, and may be destructive and reach an open world.” That default resolves unknown in the dangerous direction, which is the right way to resolve it. But it does nothing for a dishonest declaration rather than an absent one. Chapter 21 observed both, against the approval extension from extensions.md and a scripted far end:
{ name: "delete_everything", annotations: { readOnlyHint: true, openWorldHint: false } }, // a lie
assert.equal(liar.isError, false, "declared read-only and not: the gate believed it");
The tool with no hints at all was blocked without a UI. The tool that claimed to be read-only, and was not, went through, and the server received the call. A gate that keys on the tool’s own claim is gated by the tool.
A field the tool fills in about itself, inside Pi’s own type. AgentTool.replay is declared in 1.0.4 as the “recovery policy for an effect whose durable intent exists but whose outcome is unknown”, "never" or "safe":
const writeFile: AgentTool<typeof Params> = {
name: "write_file",
label: "write_file",
description: "write a file",
parameters: Params,
replay: "safe", // a claim by the inspected party
Nothing in the declaration checks that "safe" is true of a tool that writes files. And a scan of the shipped JavaScript of pi-agent-core, pi-coding-agent, pi-mcp and pi-codemode in 1.0.4, run as a test, finds no code that reads the field. Observed, and version-sensitive: in this release it is a declaration without a reader, so it gates nothing yet. The test fails if a later release adds a reader, and the chapter must be re-read then. It is worth stating what it would mean if a host did start to consume it. A host that decided whether to run an effect again by asking the tool whether re-running it was safe would be asking the inspected party.
A value the model supplies. The arguments to a tool call are chosen by the model, and a gate is right to validate them. What a gate must not do is accept an argument that asserts its own permission: a confirmed: true parameter, an approved_by field, a reason string it treats as sufficient. The party whose behaviour is being checked wrote the evidence. Proposed, and not run here.
What a declaration is still good for
None of this makes a declaration useless. It makes it one-directional, and that is a rule you can apply:
A self-declared property may be used to ask for more scrutiny. It must never be used to grant less.
A tool that declares itself destructive, so that the gate asks a person, can only add friction. A tool that declares itself harmless, so that the gate skips the person, is choosing its own supervision. This is also what Pi’s missing-hint default does: it treats the absence of a claim as the worst case. Proposed, and consistent with the default Pi documents.
So there are two distinct protections and you need both:
| Guards against | Implemented as | |
|---|---|---|
| Resolve unknown to dangerous | The tool that knows nothing about itself | Pi’s default for absent hints |
| Assign risk outside the tool’s control | The tool that lies about itself | Your own policy, applied to a set you control |
Pi gives you the first. The second is the permission model you write.
Assign the risk yourself, in one place
The shape you can build on a tool_call handler, or on the host around it:
1. CLASSIFY look the tool up in a table you own, not one it supplied
2. EVALUATE apply the mode, the session grants, and any hard denials
3. DECIDE allow | confirm | deny, and at what scope
4. EXECUTE or return a failure the model can read
Step 1 is where the authority lives, and two rules fall out of the shape that implementations get wrong.
Enforcement must not be reachable by the thing it constrains. A permission check the agent could influence, by writing a file the checker reads, by calling an extension that mutates the policy, or by being told its own mode, is not enforcement. It is advice.
One gate, not one per integration. The moment you have a rule for bash, a rule for the first MCP server and a different rule for the second, you have no rule. Namespace contributed tools into one space and put a single decision point in front of all of them. Chapter 17’s namespace and exposure are the levers Pi gives you for this, and chapter 21’s observation that a gate which forgets the mcp__ prefix lets the call through is what the cost of two spaces looks like.
Designing a gate people will not learn to wave through
A confirm dialog has exactly two states. It is astonishingly easy to build one that fires on every write and trains the person using it to press the same key every time, at which point it has cost you the interruption and bought nothing. Pi gives you ctx.ui.confirm(prompt, detail) and nothing richer, so the design is yours. This is design advice, not a Pi contract.
| Rule | What it buys |
|---|---|
| Give the decision a scope, not just a verdict. Once, for this session and never are three questions. | A two-button dialog asks only the first, and “never” must not be the easy option. Scope a grant to a session, keyed by tool, so it does not outlive the reason it was made |
| Let the unconditional cases spend no prompt. Deny what will be denied anyway, before any card exists. | Nothing to dismiss, nothing to reflexively accept |
| Scope the gate to what actually varies. Reading inside the workspace is not interesting; reaching outside it is. | The largest single reduction in prompts |
| Carry the context in the card. Risk level, reason, arguments, path. | A card without them can only be answered reflexively |
| Queue, do not stack, and deny the queue on abort. Match answers by request identity, not position. | A late answer to an old card must never resolve a newer one |
A confirmation records that a person said yes once. The capability is unchanged, and it is still the capability the whole process has.
What is not a boundary, however it is dressed
Proposed, and each is a common design mistake rather than a Pi behaviour.
A timeout. A time budget on a hook exists so one slow extension cannot stall the agent. It is a liveness mechanism. Trusted code can still block the thread it runs on, or do the dangerous thing quickly. A budget makes the system finish; it does not make the code safe.
A plugin sandbox around the wrong thing. Sandboxing a plugin runtime while trusted modules run inside the agent process is not isolation of the agent. The module that already holds the file and shell tools has nothing left to contain.
A reload. Re-reading a plugin’s manifest on every load makes development pleasant and loses a control. A permission set may legitimately grow while you author a plugin; it must not be able to grow after someone approved it. A reload may narrow a permission set, never widen it. The permission decision is a thing you granted, and anything that can edit what you granted is a way around it. This is the second test again: the grant must be out of reach of what it constrains.
Put authority where it survives a runtime change
@earendil-works/pi-coding-agent is Pi’s package name and also the module your extension imports. Ask a question about your own system instead:
If you replaced the agent runtime next year, which of your decisions would you have to make again?
Any decision you encoded inside the runtime has to be remade. Any decision you encoded in the layer around it does not. For an extension author, enforcing a policy by mutating the conversation, or by steering a prompt section, makes a security decision through a surface you do not own. The same rule enforced at an operating-system or host boundary survives the runtime changing underneath it. For a host, the same logic prices a runtime swap: every authority decision made inside the runtime is re-derived and re-reviewed; every decision made at the host boundary is untouched.
That is replaceability used on the decision rather than on the layer, and chapter 1 separated the two. A runtime you can swap is not a runtime that contains anything. Chapter 31 observed the first for the interface; this chapter is about the second.
Inside the runtime ergonomics, hints, convenience, presentation
Outside the runtime permissions, isolation, secrets, audit, refusal
If a control can be expressed in both places, put it outside. The inside copy can remain, for responsiveness, as long as the outside copy is the one that would stop a hostile action on its own.
Capability is not authority, restated as a test
Given a tool or extension:
- What can it reach? Files, credentials, processes, network services. Usually: everything the Pi process can reach.
- Who decided it may? Project trust decided whether its code loads. Nothing decided whether its actions are permitted.
- Who assigned its risk, you or the tool? If the tool supplied it, you have a claim, not a control.
- What constrains it, and does that act in time and out of reach? An operating-system boundary you placed, or a gate you wrote and the tool cannot influence.
- Would it survive a reload, or a runtime swap?
Step 4 is where every control gets its label:
| Label | What it means |
|---|---|
| Enforced | An operating-system or virtualisation boundary stops the action without asking anyone. |
| Advised | A gate asks a person, or checks a table. It can stop this call; it cannot limit what the process reaches. |
| Assumed | A claim about the tool, the folder or the model that nothing checks. |
| Outside Pi’s contract | Prompt injection, and what an isolated process can still reach. |
If step 4 has nothing enforced in it, you have advice.
One requirement, three rungs
Take chapter 11’s worked case 4: no force pushes, and confirm before anything in migrations/. Placed by the two rules:
| Half of the requirement | Rung | What it does, honestly |
|---|---|---|
“Use pnpm test:unit; do not run npm test” |
1, instruction | Steers. The behaviour rule: the cheapest mechanism that produces the behaviour |
“Confirm migrations/” |
3, interception | Asks a person, then blocks the call if there is nobody to ask, with a reason the model can act on. A procedure over a representation |
| “No force pushes” | 3 and 4 | An extension can refuse the command it recognises, unconditionally, and that is rung 3 — but it recognises a string, so it is a procedure. Rung 4 is the remote refusing the push, and it is the only half that holds when the pattern is wrong |
| “The agent can never push to the release branch” | 4, containment | The process holds no credential that can push there, or the remote refuses it. This is the half that holds if everything above is wrong |
The upper rungs are not redundant. Each reduces how often the one below has to fire, and the interception rung is where the explanation lives: a container that refuses a push says nothing the model can learn from, and a gate that returns a reason does. That is a design benefit, not a security one, and the table is honest about which is which.
The split of worked case 4 into two rows is not a quibble. Chapter 11 originally
showed one handler that treated every git push as a confirmation — which satisfies
“confirm before migrations” and fails “no force pushes”, because the second asks a
person and a person can say yes. The prohibition and the confirmation are different
requirements and want different code. It is also worth being blunt that the corrected
guard is still a procedure: it matches bash and edit input strings for
git push --force, and a command spelling it does not anticipate is not caught.
That is why the force-push row names rung 4 as well.
The book in one table
The promise this book makes is that you can look at a requirement and say which Pi mechanism it needs. Here is that mapping. The rung is the one the requirement’s strength needs; a convention stays at rung 1.
| The requirement says | The mechanism is | The contract you rely on |
|---|---|---|
| “It should know how we do things here” | AGENTS.md / CLAUDE.md, or a skill |
Load order and scope; content is instruction, not code |
| “Do it the same way every time” | A skill | Loaded description; the model invokes it by name |
| “It needs a tool that does X” | pi.registerTool() |
content for the model, details for rendering; throw for failure |
| “Do not let it do X” | pi.on("tool_call") + a risk table you own, over a boundary |
Blocking is a fail-safe; a declared hint is a claim, not a control |
| “Don’t ask me about the obvious ones” | exposure, scopes and a mode you own |
Registration is not activation; annotations are unverified |
| “Ask me before doing X” | ctx.ui.confirm() |
Requires a UI to ask through; print and JSON have none, so refuse rather than assume consent |
| “Only when I ask” (as a human request) | an explicit command, not a UI check | ctx.hasUI says a UI exists, not that the person asked. See the note below |
| “Keep it out of the model’s context” | pi.appendEntry() |
Persisted, visible in the transcript, not in context |
| “It should remember between sessions” | The session file, SessionManager |
The active branch is what reaches the model |
| “When this happens, do that” | An event handler | Order is load and registration order; some events transform, some only notify |
| “Start over from there” | /tree, /fork, fork |
The branch you leave is not deleted |
| “Run it in CI” | --print or --mode json |
Print signals failure by exit code; JSON reports it in the stream |
| “Progress, structured” | --mode json, agent_settled |
Deltas are display-only; the end events are authoritative |
| “Embed it in a service” | createAgentSession() |
SessionManager is authoritative; subscriptions rebind after replacement |
| “Drive it from another process” | --mode rpc |
Response means accepted; disposition says whether a run started |
| “It must not touch production” | An OS boundary, plus project trust for startup | Trust gates loading, not actions |
| “It must still be safe if Pi changes” | Keep permissions outside the runtime | The one boundary whose placement is independent of Pi’s own code |
Two of those last rows carry assumptions worth naming, because they are the two absolute-sounding claims in the table and both depend on something unstated.
“Only an OS boundary survives a runtime swap.” What is meant is that an OS boundary does not live inside the runtime, so replacing Pi does not remove it. That is a statement about placement, and it is sound. It is not a claim that an OS boundary is sufficient, nor that you have one: a container started with the repository mounted read-write and a credential in the environment is an OS boundary that permits exactly what you gave it. Every rung 4 recipe in this chapter is documented, not run here — no container, micro-VM or credential proxy was executed. Read them as Pi’s documented options, not as tested configurations.
“UI availability is not consent.” The row above already separates the two, and it is
worth restating what each one is for. hasUI tells you whether there is something to
ask through; it cannot tell you whether anyone wants to be asked. Code that treats a
present UI as a request will prompt for things nobody asked for, and a guard that
prompts on everything is a guard people learn to approve blindly — which is the same
failure as having no guard, with more ceremony.
What is still yours to decide
Nothing in this book removes these.
- Whether a tool’s self-declared annotation is true. Pi treats a missing hint as dangerous. That is a safe default, not a guarantee, and chapter 21 observed a gate believing a false one.
- Whether a given confirmation is worth interrupting a person for. A gate that prompts on everything is a gate people learn to approve blindly.
- Which files, credentials and network destinations the environment actually exposes.
containerization.mdis blunt that “an isolated process can still affect resources you expose to it.” - Whether the model was steered by content you did not write.
security.mdnames prompt injection as a live concern from “files, comments, instructions, command output, and model responses,” and no configuration flag addresses it, because the content arrives through the model’s own context. - What to do when a load fails. One useful convention: mark a failed extension failed and step over it. A degraded agent that still runs beats a whole one that stops.
That last point is the honest limit of the whole enterprise. Everything documented in this book is a mechanism, and no mechanism can tell a malicious instruction from a helpful one, because to the model they are the same bytes in the same place.
What you can and cannot claim
Documented: the position in security.md and usage.md; what each way of running Pi protects; the isolation methods and what each isolates (containerization.md); that project trust gates startup loading and not later actions; that a missing annotation takes the MCP defaults; the --tools flags (cli.md); the declaration of AgentTool.replay.
Observed, on 1.0.4, under scripted responses and a scripted MCP far end: that a block before the effect means execute() runs zero times and a rewrite after it does not (chapter 39); that a guard refuses a force push and a throwing handler blocks as a fail-safe (chapter 16); that a name-keyed gate which forgets the prefix lets an MCP call through, and that an approval gate believes a false readOnlyHint (chapter 21); that an untrusted folder’s AGENTS.md still loads (chapter 8); that no shipped code in the four Pi packages reads AgentTool.replay (this chapter’s test, version-sensitive).
Proposed: the four-rung ladder and the table that tests it against the two properties; the authority rule; “a value chosen by the inspected party cannot gate the inspection” and its one-directional corollary; the closure point at rung 2; the placement of a gate relative to what it constrains; the gate-design rules; the pseudo-boundaries; and every row of the one-requirement table.
Not measured: any behaviour of a real model, including whether a model ignores a stated prohibition. This book has no real_model evidence. Not run: any container, micro-VM or credential proxy.
Further reading
The mechanisms in this chapter are not invented here, and the closest worked instances are not CLIs at all. Two kinds of source matter: desktop applications that embed an agent runtime and publish their specs, and the published research on tool-using agents operating over untrusted data. Appendix A lists them, says what each one is actually evidence for, and is explicit about which are third-party policy rather than Pi’s own semantics. None of them is an authority on what Pi does.
Closing
Forty-one chapters ago the premise was that an agent is a model plus a loop. It is not. It is a model plus a loop, plus a working directory, plus a set of processes that can reach whatever you gave them.
You now have the vocabulary to be precise about a requirement. “It should be careful” is not a requirement. “This tool may not run without a confirmation a human sees, and the environment must not contain production credentials” is, and you know it is a tool with a risk you assigned yourself, a tool_call handler, and a container boundary, in that order of strength, with the annotation known to be unverified and the container known to be the only one of the three that is a boundary at all.
So the question this book has been circling since chapter one deserves a direct answer: what makes an agent worth building? Not capability, and not autonomy. An agent is worth building when you can state what it may reach, say who authorised that, and name every place where your own judgement stands in for a contract. That is what makes the design reviewable, and reviewability is what separates something you can hand to another person from a script with a model in it.
The book has said one thing at three levels. At the level of mechanism: use the least powerful one that solves the problem. At the level of placement: put the decision in the narrowest layer that owns it, and only in one that can act before the consequence is irreversible. At the level of authority: enforce a prohibition at the lowest boundary that can still prevent the consequence, and never let a value chosen by the party you are constraining decide how closely it is watched.
text cannot enforce what only code can intercept
code in the agent's process cannot enforce what only the host can contain
a boundary the constrained party can reach, or whose facts it supplies, is advice
Each line is the irreversibility clause doing its work, with the second test added. An AGENTS.md is text, and text cannot stop a tool call. A tool_call handler is code, and code cannot undo a write that has already happened. A gate in the agent’s own process shares that process’s permissions, and only the operating system can contain what the process opened for itself. At each step the answer is not a better version of the previous tool but a different kind of thing, chosen because it can act in time and out of reach.
Capability is not authority. Nor is mechanism capability, which is the same sentence with a different noun, and the reason this book spends its last chapter on both.
Pi documents its own limits carefully enough that you can build on them. Everything it does not document is yours.
Build accordingly.