From Agent Architecture to Agent Engineering
A retrospective on what happened when this book’s architectural ideas were implemented, instrumented, challenged, and measured.
What This Chapter Is
The first edition of this book proposed an architecture: roles, tools, memory, reflection, coordination, versioning, human direction.
After it was written, those ideas were built — context systems, memory systems, evaluators, handoffs, agent runtimes — and the building forced several of them to become more precise. Some survived contact with implementation. Some failed. Some turned out to be questions about measurement rather than questions about agents.
This chapter is that retrospective. It is organized around three evidence chains, each with the same shape:
original principle
↓
implementation
↓
observed failure or limitation
↓
measurement
↓
refinement
↓
revised principle
Nothing here is a survey of current technology. The experiments are small, fixture-bound, and reported at exactly the strength they earn. Their value is not scale. It is that they show what the architecture looks like after it has been observed.
Chain 1 — Context: Retrieval Is Not Delivery
The original intuition was simple: an agent needs relevant context, so retrieve it.
Implementation forced the idea apart. Between “the information exists” and “the model used it” sits a chain of distinct transitions:
information exists
↓
information is eligible
↓
retrieval
↓
selection
↓
exposure
↓
possible influence
↓
possible utility
Each link needs its own mechanism, and confusing any two is one of the most common errors in agent design. A stored record must be retrieved, then selected, then actually assembled into the context bundle — and even exposure does not guarantee influence on the output, let alone useful influence.
The test was deliberately narrow. Twelve small tasks, each with one decisive span and three distractors, assembled into a context bundle under different selection policies, with a byte-level check of what the bundle actually contained. No model call was involved: this run tests delivery to the bundle, not influence on output.
The first run, under a generous budget, did not distinguish the policies: everything fit, so everything was exposed. That non-result was kept, not discarded — and a withholding control, which removed the decisive span before assembly and exposed it in none of the runs, showed that the exposure instrument could register non-exposure.
The second run used a tighter budget, repeated across three seeds. The 60-unit ceiling was chosen after an 80-unit trial proved loose enough for everything to fit, and the protocol was frozen before the run:
retrieval-order assembly: exposed the decisive span in 36/36 runs
gated assembly: exposed the decisive span in 36/36 runs
distractor-first assembly: exposed the decisive span in 0/36 runs
withheld control: exposed the decisive span in 0/36 runs
Retrieval occurred in every arm. Exposure did not. Under constrained budgets, ordering policy determined whether the decisive information was in the bundle at all. Retrieval order was fixed by construction, decisive span first, so the result follows from the budget arithmetic: it demonstrates a failure mode and the instrument that detects it, not how often real systems fail this way.
The revised principle:
A context system should not claim success merely because information was retrieved. It needs evidence about what was actually selected and delivered to the model.
That is the conceptual move from retrieval architecture to transport and exposure architecture. It does not claim delivered context influenced any answer — that would need a separately measured link. It claims something prior and load-bearing: without delivery evidence, influence claims have nowhere to stand.
Chain 2 — Memory: Storage and Retrieval Are Not Enough
The original book treated memory well but incompletely: store information outside the model, retrieve selected pieces into context later, add eligibility, provenance, supersession, and conflict rules.
The engineering problem is what happens when retrieval works exactly as designed and returns the wrong past:
relevant memory
stale memory
superseded memory
conflicting memory
unsupported memory
All of it retrievable. All of it capable of entering context with equal confidence.
The test loaded eight small question-answering stores with exactly that mixture and compared raw retrieval against an answer-blind admission gate — admit current notes, flag conflicting ones as disputed, drop superseded, stale, unsupported, and irrelevant material:
raw retrieval: 3/8 correct
admission-gated: 5/8 correct
good notes rejected by the gate: 0
Two mechanism cases show what the gate did. Dropping a superseded note (“check window seals”, replaced weeks ago) recovered the correct “door seal” answer. Flagging a disputed scheduling claim as contested let the model choose the current date instead.
The third case matters more. In one task, the gate flagged a conflicting figure as disputed — and the model selected it anyway. Flagging alone was insufficient. That counterexample stays in the record because it bounds the claim: admission changed outcomes on these fixtures; it did not solve contamination. Two further caveats apply. The gate read each note’s ground-truth tag from the fixture, and the fixtures were built to contain exactly the failure classes listed above, so this demonstrates mechanisms rather than discovering them, and is not a realistic admission benchmark. And two answers that were semantically right were scored wrong by the substring checker in both arms.
The revised principle:
Memory architecture requires admission, provenance, supersession, and conflict handling in addition to retrieval.
Or compactly:
memory ≠ database + search
memory =
storage
+ retrieval
+ admission
+ exposure
+ lifecycle
+ conflict handling
No general memory-performance claim is made. Eight fixtures and one small model (qwen2.5:0.5b) cannot carry one. The lesson is architectural: storage is a decision about what may influence the model, and it needs machinery of its own.
Chain 3 — Evaluation: The Measurement System Is Part of the System
This is the strongest chain, and the most uncomfortable one — because the failure it exposes is in the evaluator, not the agent.
The familiar assumption: evaluate the agent’s output, compare single, single-with-critique, and three-role organizations under an equal call budget, and read off which architecture wins.
The first real-model attempt (qwen2.5:0.5b) returned a floor: zero fully-verified successes in all three arms. The second attempt, with a larger model (gemma3, 4.3 billion parameters) and retained raw outputs, returned the same floor. At first glance, a model limitation.
Then the second run’s retained outputs were read. That model had been extracting the intended values all along:
model wrote: 200 checker expected: "200 euro"
model wrote: Router kestrel-9 checker expected: "kestrel-9"
model wrote: -18C checker expected: "minus 18C"
Content correct; representation rejected. The diagnosis moved from agent failure to what was substantially an evaluation failure: a brittle exact-string checker had manufactured an apparent floor.
That changed what the third round was for. Its purpose was not to obtain a preferred architecture result but to repair the evaluator. Semantic field rules — numeric equality for currency, unit-aware comparison for temperature, code-with-optional-prefix for identifiers, bounded substring for short text — were frozen in writing before execution — but only after the failing outputs had been read, so they target mismatches already seen, and the new checker was never validated against a human-labelled set of equivalent and non-equivalent answers. Both scorers were kept side by side so the effect of evaluator choice would stay visible:
strict exact-match checker: 0/6, 0/6, 0/6
semantic field checker: 5/6, 5/6, 5/6
Same model. Same prompts. Same task family. The headline moved from zero to five of six because the measurement changed. That establishes that evaluator design materially changed the measured outcome here; it does not establish that the semantic checker is ground truth.
And the architecture comparison itself returned a null: single, critique, and multi-role organizations were indistinguishable at 5/6 each on these six tasks. That null is reported, not rescued. Six tasks cannot rank architectures, and no winner was hunted.
The revised principle:
Agent evaluation is itself an engineered subsystem.
A score is produced by a checker with its own strictness, blind spots, and format assumptions — the same scrutiny the book applies to critics, reviewers, and judges applies to the harness that judges them. This connects directly to Chapter 5: a judge must earn trust, false accepts and false rejects must be reported separately, and a verdict of “0/6” says as much about the evaluator as about the evaluated until the evaluator has been inspected.
The Bigger Transformation
The early architecture of this book could be summarized as:
model
+ role
+ tools
+ memory
+ reflection
+ coordination
The engineering view earned through the work above is longer:
intent
context assembly
model proposal
authority
execution
observation
state
memory lifecycle
evaluation
adjudication
trace
rollback
human control
Or as one loop:
intent
↓
context
↓
proposal
↓
authority
↓
execution
↓
observation
↓
evaluation
↓
state / memory
↓
next action
This is not presented as a universal canonical architecture. It is the architectural vocabulary this book has arrived at: each line names a transition that the book argues needs its own mechanism, and several of them — context delivery, memory admission, evaluation — were implemented, instrumented, and measured here on small fixtures. The model proposes; the runtime constrains; the environment responds; the observer records; the evaluator judges; the authority decides.
Observability: What We Cannot See, We Cannot Claim
One lesson runs beneath all three chains, and it needs no fourth experiment: if the relevant transition was not observed, the result cannot be confidently explained.
The distinctions are already in the book:
trace ≠ explanation (a model's story about its work is not a record of it)
retrieval event ≠ exposure (finding is not delivering)
tool proposal ≠ execution (suggesting is not doing)
successful execution ≠ correct outcome (running is not succeeding)
That is why logs, receipts, event records, and frozen artifacts recur throughout these chapters. They are not bureaucracy. They are the only way to tell, afterward, which part of the system did what — and the architecture-comparison story shows what happens without them: a floor with no visible cause, interpretable as anything until the raw outputs were retained and read.
Failure as Evidence
The failed and inconclusive runs in this book are load-bearing, not decorative. Their shape is worth naming because it is the method:
valid experiment
→
unexpected null
→
instrument inspection
→
measurement defect found
→
protocol revision
→
new result
→
original architecture question still null
The architecture-comparison experiment followed that shape exactly — and the last line is the point. The protocol got better; the architecture question stayed open. A failed claim used this way is more useful than a successful demonstration, because it identifies the part of the system that was not yet understood. In this case the misunderstood part was the checker, and “fix the evaluator” turned out to be the engineering.
This is a methodological lesson, not a motivational one. Failure is not celebrated here. It is instrumented, bounded, retained, and allowed to redirect the work.
What Survived
The original ideas that still stand are the spine of the first edition, intact:
model ≠ agent
roles matter as bounded responsibilities
tools require controlled boundaries
memory requires architecture
reflection requires evaluation
coordination requires interfaces
human intention remains the anchor
What changed is precision:
prompt → context bundle
memory retrieval → memory lifecycle
reflection → candidate + evaluation + adjudication
tool use → proposal + authority + execution
multi-agent teamwork → responsibility + interface + cost
success → measured outcome under an explicit evaluator
The first book said what the components are. The revision says how each component earns its claim — and what to check when the claim fails.
Where This Leaves the Work
Agent architecture begins with deciding what components exist.
Agent engineering begins when we can observe how information moves, who has authority, what actually executed, how outcomes were evaluated, and whether the system preserved the intention that started the process.
Those are observable properties of a running system, not attributes of a model. They are what the three chains above learned to check — delivery rather than retrieval, admission rather than storage, evaluator behavior rather than headline scores.
The question of what all this capability is for belongs to the chapter before this one. This chapter answers the narrower question underneath it: how we know the architecture actually worked. The agency the book ends with is worth more for being checkable.