Appendix A — References and Further Reading
The sources behind this book: Pi's own documentation, published research on agent loops, interfaces, context and security, ecosystem projects that demonstrate particular designs, and the companion books that take these questions further.
This book makes three kinds of statement and keeps them apart: what Pi’s documentation and exported types say (documented), what was run and recorded (observed), and what this book recommends (proposed). A reader who wants to check a claim needs to know which pile it came from, so this page says where the material lives.
Nothing below is evidence about how Pi behaves unless it is Pi’s own documentation or source. External work is here because it illuminates the same problems, not because it defines Pi.
How to read a citation here
The three source kinds below behave differently as evidence, and mixing them up is how a book’s bibliography turns into its weakest chapter. This is the same distinction the prose uses, applied to the sources themselves.
flowchart TD
Q{"what is the claim about?"}
Q -->|"Pi's behaviour"| D["docs/*.md and dist/*.d.ts<br/>at the pinned release"]
Q -->|"a general result<br/>about agents"| R["a paper<br/>read the full text, not the abstract"]
Q -->|"a design someone<br/>implemented"| E["a repository<br/>proves it exists, not that it is better"]
D --> D2["assert it plainly, cite the topic<br/>declarations beat prose"]
R --> R2["attribute the finding to that paper<br/>numbers belong to their experiment"]
E --> E2["mark it proposed<br/>a fork's reasons describe the Pi it met"]
D2 -.->|"the rule that governs all three"| L["check the 1.0.4 pin before accepting any of them"]
R2 -.-> L
E2 -.-> L
The dashed line is the one to keep in mind. A paper or a repository read this month says nothing about Pi 1.0.4 unless it is about Pi 1.0.4, and an external comparison that cannot be pinned to a release is a proposed claim whatever the confidence in it.
Pi primary sources
Everything documented in this book is checked against the installed package for the pinned version, not against memory:
pi --version
~/.pi/agent/install/<VERSION>/node_modules/@earendil-works/pi-coding-agent/
Inside it, docs/ holds 40 topic files at 1.0.4, examples/extensions/ holds runnable
examples, and the exported declarations under dist/ are authoritative when
prose and types disagree. The upstream repository is earendil-works/pi, and
each release’s CHANGELOG.md is the first place to look when a claim stops
matching what you observe.
Agent loops and tool use
-
ReAct — Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models, ICLR 2023, arXiv:2210.03629. The original reasoning-then-acting loop. Useful here as the ancestor of the turn structure in chapter 1, and as a reminder that interleaving observation and action is the whole idea.
-
Inside the Scaffold — Inside the Scaffold: A Source-Code Taxonomy of Coding Agent Architectures, Rombaut, arXiv:2604.03515. Surveys thirteen open-source coding-agent scaffolds at pinned commit hashes, and finds eleven of thirteen compose multiple loop primitives rather than relying on a single control structure. Its context-compaction taxonomy spans seven distinct strategies, and its tool counts range from zero to thirty-seven. The load-bearing point for this book: compaction is not one mechanism, and a “narrow runtime” is one point in a design space rather than the settled answer. Unrefereed preprint; treat its specifics as directional.
-
Towards a Science of Scaling Agent Systems — arXiv:2512.08296. Ran 260 configurations across six agentic benchmarks and three LLM families, with tools, prompts and compute standardised. Relative change against a single agent ran from +80.8% on decomposable financial reasoning to −70.0% on sequential planning, and the predictive model reaches a cross-validated R² of 0.373 (0.413 with a task-grounded capability metric), identifying the best architecture for 87% of held-out configurations.
Our reading, not the paper’s: this is the strongest evidence we found that adding orchestration is a gamble, and it belongs beside any advice to reach for a more elaborate scaffold. The paper’s own conclusion is narrower and more positive — coordination pays where the communication topology matches the task structure, not because more agents is better. Preprint.
Agent–computer interfaces
-
SWE-agent — Yang et al., SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, arXiv:2405.15793, NeurIPS 2024. Treats the interface given to an agent as a design variable: with the model’s weights unchanged, a purpose-built agent-computer interface beats a default Linux-shell interface by 10.7 percentage points on SWE-bench Lite. The reason chapter 17 treats how a tool is presented as a decision rather than an afterthought.
-
CodeAct — Wang et al., Executable Code Actions Elicit Better LLM Agents, ICML 2024, arXiv:2402.01030. Consolidating many tool calls into one executable-code action space reports up to a 20-point absolute gain in success rate on M³ToolEval across 17 models, with up to 30% fewer actions.
Included as a counterexample: the book’s claim that a schema-driven tool surface is worth having is a trade-off, not a proof. Two qualifications keep the citation honest. The gain is specific to multi-tool composition — on API-Bank’s single atomic API call, CodeAct is the best format for only 8 of 17 models. And the effect is not a fine-tuning artefact: the same base models are run in all three action formats, so the comparison isolates the format. What the paper does concede runs the other way — its respectable JSON scores for closed models suggest those models were fine-tuned toward JSON, which flatters the schema-driven baseline.
Context and memory
- MemGPT — Packer et al., MemGPT: Towards LLMs as Operating Systems,
arXiv:2310.08560. Argues that long-running systems
need explicit policy for what stays in active context. Supports the general principle
behind chapter 4; it says nothing about when Pi compacts, which is Pi’s own documented
contextTokens > contextWindow − reserveTokens. - Lost in the Middle — Liu et al., Lost in the Middle: How Language Models Use Long Contexts, TACL 12 (2024), arXiv:2307.03172. Relevant to chapter 4’s warning that a large window is not the same as usable attention: accuracy degrades with position, best at the beginning and end of a long input and worst in the middle. Honest limit: it establishes no usable-context threshold, and neither should the book. (The arXiv record’s Comments field still reads “TACL 2023”; the published volume is 2024, which is what the book uses.)
- The companion books Context From First Principles and Memory From First Principles develop admission, scope, survival, pruning and retrieval in general. This book only needs Pi’s concrete terms.
Evaluation and reliability
-
τ-bench — Yao et al., τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, ICLR 2025, arXiv:2406.12045. Introduces
pass^k— the chance that all k independent trials of a task succeed, averaged over tasks. It is the mirror image ofpass@k, which asks for at least one success in k. The paper’s own numbers make the point: gpt-4o with function calling reachespass^1of roughly 61% on τ-retail and about 35% on τ-airline, yet itspass^8on τ-retail falls below 25% — on more than three tasks in four it fails to succeed on all eight attempts. A configuration averaging 61% is still unreliable nearly eight times in eight.The practical consequence for a reader of this book: a demo is a single draw, not a capability claim, and chapter 10’s advice about writing a finishable task is partly this. The earlier draft of this appendix garbled the metric by attaching the 25% figure to a single-run success rate and inverting what “unreliable eight times out of eight” means; the numbers above are the paper’s.
-
SWE-Bench Illusion — Liang, Garg & Zilouchian Moghaddam, arXiv:2506.12286. Probes model knowledge rather than evaluating agents: naming the buggy file from the issue description alone, and reproducing the ground-truth function from file context alone. State-of-the-art models reach up to 76% path accuracy from the issue text alone against up to 53% on repositories absent from SWE-bench, and verbatim function reproduction reaches up to 35% consecutive 5-gram accuracy on SWE-bench Verified and Full against up to 18% elsewhere. The paper’s own conclusion is deliberately hedged — gains “may be partially driven by memorization rather than genuine problem-solving” — and it argues for contamination-resistant benchmarks rather than declaring the leaderboards invalid.
-
Does SWE-bench-Verified Test Agent Ability or Model Memory? — Prathifkumar, Mathews & Nagappan, arXiv:2512.10218. A separate study, with separate comparators, and it is the one that uses logically impossible tasks: it tests two Claude models on file localisation only, withholding repository context “to the extent the task should be logically impossible to solve”. Against comparable Python projects in BeetleBox and SWE-rebench, the models score roughly 3× higher on SWE-bench-Verified, and 6× higher at naming edited files with no project context at all — which it argues biases leaderboards toward agents that happen to use those models rather than toward strong agent design.
Two papers, two findings, attributed separately on purpose. An earlier draft merged them into one bullet, which borrowed the memorisation framing for one and the impossible-task framing for the other. Any leaderboard number quoted alongside this book should carry the caveat, and the caveat belongs to a named paper.
Security, prompt injection and authority
- AgentDojo — Debenedetti et al., AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents, NeurIPS 2024 Datasets & Benchmarks Track, arXiv:2406.13352. A tool-using agent operating over untrusted data is a distinct security object, and the attack surface arrives through tool output rather than only through user input. Its results also show how weak the baseline is before any attack: the best model tested, Claude 3.5 Sonnet, solves only 78.2% of benign tasks with no attack present at all, and GPT-4o 69.0%. A benchmark where the undefended system already fails roughly a quarter of ordinary tasks is why chapter 42 refuses to treat sanitisation as containment.
- Pi’s own
docs/security.mdis the authority for Pi’s position, and chapter 42 quotes it directly. Third-party permission systems exist in the ecosystem and are good implementations of a decision layer; they are not Pi’s semantics and are not presented as such.
Agent runtimes and architecture
For the layering argument in chapters 1 and 11, the primary evidence is Pi’s own package graph. One number here is verified and the rest is not, so they are separated.
Verified. pi-agent-core at 1.0.4 imports @earendil-works/pi-ai, typebox,
and relative modules only. Its two declared runtime dependencies are those first two,
and every module specifier appearing in its dist/ is one of them or a relative path:
zero interface imports. The transitive graph closes the question too — pi-ai
depends on provider SDKs, telemetry and schema libraries, and on no UI or TUI package.
The number is worth stating precisely because it is countable, and it is stated at
1.0.4 so that it cannot drift silently under a reader who is looking at a later
release.
Not pinned. The four comparisons below are proposed — hedged reading, not a measured dependency graph. No release or commit for any of them is pinned anywhere in this repository, which means a reader cannot reproduce them and a release can move them. Treat the direction of each as the claim and not the detail.
- pydantic-ai — appears to achieve an embeddable runtime by splitting
distributions, with per-provider extras on a slim package. Reading this release’s
packaging suggests the boundary is declared rather than conventional, but no
pyproject.tomlwas read at a pinned version, so that is not established. - smolagents — has a hard, non-optional
richdependency and UI modules in the same source tree as the agent code, which is the shape chapter 11 calls boundary dissolution. Directionally supported at the versions inspected; not pinned. - opencode — note the name is ambiguous: it is the coding-agent repository
anomalyco/opencode and a Pi provider id
reached through the OpenAI-responses API. The repository’s
packages/layout is genuinely layered, and its LLM package genuinely isolated, which makes it the closest structural analogue to Pi. The claim that its agent loop is not a library primitive is not established, and one datapoint cuts against it: the published npm package declares zero runtime dependencies and ships a single prebuilt executable, which is bundling rather than process separation. The contrast this paragraph sets up — that Pi achieves the same separation by import rather than by process — is therefore unresolved, and the verified zero-import number above should not lean on it. - LangGraph — is an embeddable runtime, and its agent behaviour lives either in user code or in a prebuilt package that has been moved out of its own dependency graph. It is the least pinnable of the four: hundreds of published versions and a release within days of this writing.
Pi ecosystem projects
A repository’s existence proves someone implemented a design; it does not prove the design is better. These are included because each demonstrates one idea clearly, not because of popularity.
Every project below is reachable, and none was pinned to a commit for this edition. A repository’s existence proves someone implemented a design; it does not prove the design is better, and an unpinned HEAD can change what the code does.
| Project | What it demonstrates |
|---|---|
earendil-works/pi-review |
A complete review workflow built entirely above the runtime: zero provider changes, zero core changes, zero new tools. |
earendil-works/gondolin |
A real isolation boundary. Its Pi example works because Pi exports tool operations separately from tool definitions, so built-in tools can be redirected into a micro-VM while the agent process stays outside. |
nicobailon/pi-boomerang |
Context collapse as a first-class session-tree operation rather than something re-derived by a plugin. |
nicobailon/pi-memory-workbench |
A memory system with an explicit invariant set, markdown as source of truth, and honest documented limitations. |
MasuRii/pi-permission-system |
A permission decision layer on Pi’s public hooks, with a threat model that explicitly refuses to call itself a sandbox. Third-party semantics, not Pi’s. |
svkozak/pi-acp |
One runtime driven by another UI protocol, over --mode rpc, with documented gaps. |
mitsuhiko/agent-stuff |
Subagents that stay outside the loop: it spawns real child pi processes rather than reimplementing the loop. |
code-yeongyu/senpi |
The boundary test. An extension-first fork that records, per change, why an extension could not handle it — the clearest available account of where Pi’s seams are. Read it as proposed evidence: a fork’s reasons describe the Pi it was written against, and Pi 1.0.4 has prompt sections, compaction hooks and overflow recovery that a fork’s history may predate. |
can1357/oh-my-pi |
The counterweight: what a distribution does when a minimal harness absorbs native FFI, telemetry and in-runtime compaction. |
thevibeworks/awesome-pi-agent |
A discovery index for the ecosystem. Useful as a list; not an authority on anything. |
Comparative agent frameworks
pydantic-ai, smolagents, langgraph, mastra and opencode, read on the
axes in chapter 1: where the provider boundary sits, where the loop lives, where
state lives, and whether the runtime embeds independently of a UI.
Two of the five appear to achieve a cleanly embeddable runtime, and one of those by
never having had a UI to separate from. Placement of state seems to distinguish
them more than placement of the loop: most put persistence inside or beside the loop,
while Pi’s runtime holds no session at all. This paragraph is proposed and is
weaker than it looks — no comparison release is pinned for any of the five, so treat
the direction as the claim and do not quote a count as a finding. The one thing here
that is measured is Pi’s own side: zero interface imports at pi-agent-core 1.0.4.
Companion books
| If you want | Read |
|---|---|
| Why a loop, in general | Agents From First Principles |
| Orchestration, subagents, multi-agent patterns | Advanced Agents From First Principles, Agent Architectures |
| What goes into context, and what is admitted | Context From First Principles |
| What persists, is retrieved, or is merely stored | Memory From First Principles |
| Diagnosing failure, state and observation | Debugging AI |
| Why a model behaves as it does | Models From First Principles, Hallucination From First Principles |
| Engineered systems rather than models | Applied AI |
Applied AI is worth singling out: it argues for separating the runtime from the model so that models can be swapped, which is the substitutability argument. The argument this book makes is about layer ownership — which layer owns a decision — which is a different question and does not restate that one.