Chapter 05 of 30

Meat Proxy

Concepts

CHAPTER 05 โ€” Meat Proxy

STATUS

Full first draft

EDITORIAL PASS (2026-09-14)

  • Visual-mechanism pass 2026-09-14: added meat-proxy feedback-loop flowchart (success, trust, thinning checks, unexamined failure, evidential decay, blast radius, with undetected-failure feedback edge). All prose claims pre-existing; diagram verified to render.
  • SOURCES VERIFIED (WebFetch 2026-09-14):
    • Autor et al., NBER w35720 (Sep 2026): pre-registered 3-month RCT, 133 patent lawyers, 11 US firms; +0.34 SD (10d), +0.38 SD (90d); unaided redline +0.32 SD overall, +0.45 SD seniors; juniors no average gain, bifurcated; Google funded direct costs. ATTRIBUTION FIXED: “foundational expertise may be a prerequisite …” is the authors’ own abstract sentence, not the book’s interpretation; only the “one-way door realest where expertise was thinnest” extension is the book’s.
    • Huang et al., arXiv:2512.14012 (v1 Dec 2025, rev Aug 2026): field observations N=13, surveys N=99; experienced developers retain agency in design/implementation.
    • Anthropic, “Agentic coding and persistent returns to expertise” (16 Jun 2026): ~400,000 sessions (~235,000 users); users ~70% planning, Claude ~80% execution; verified success ~15% novice vs 28-33% intermediate/expert; caveat: transcript analysis, not outcome validation.
  • Stale chapter refs fixed (pre-renumbering): ledger/checks -> Ch14 & 16; preserved raw output -> Ch11 & 17; claims with evidence -> Ch18; checks ahead of attention -> Ch3, 14, 21; review record -> Ch14 & 18. The ENGINEERED-REVIEW MECHANISMS chapter list below is likewise stale.
  • Added: mechanical checks exact but not sufficient (Ch21 A05); later verifiers held to seeded-corruption standard.
  • Spelling: favour -> favor.
  • Score ~895 -> ~951.

CENTRAL QUESTION

What has to be true of a process for a human review to still mean something after the hundredth time?

SECTION OUTLINE

  • Ten weeks: the drift from careful review to signature, shown step by step.
  • The inversion: it is caused by the system’s successes, not its failures. Every step is locally rational.
  • Name it โ€” automation bias (omission/commission) and automation complacency โ€” and give the 30-year evidence.
  • Measured in this setting: confidence in AI vs self-confidence; the shift toward verification/stewardship.
  • Why the model cannot supply the missing part; the definition of slop.
  • Why the failure is catastrophic rather than gradual: deskilling, then the missing trail / unbounded blast radius.
  • The institutional trap: oversight requirements as accountability shield.
  • “If you are going to do it anyway” โ€” serious guidelines for the shallow path, offered without sarcasm.
  • The real path: review is architecture, not willpower. Six mechanisms, all built later in the book.
  • The heart at the end of the machine.

LOAD-BEARING CLAIMS

  1. The meat proxy is produced by success. A 2%-error system is more dangerous to review quality than a 30%-error one, because it removes the reason to look.
  2. Diligence is not a plan. Complacency appears in experts and is not overcome by practice.
  3. Slop = locally plausible, globally unaccountable. You cannot detect it by hunting for mistakes; you detect it by asking what it claims and what would falsify the claim.
  4. The catastrophic failure is not a bad output. It is an unbounded blast radius โ€” a process that approved everything and retained evidence about nothing.
  5. Review must be engineered so the careful path is the cheap path. Willpower interventions are unsupported by the evidence.
  6. This is the reader’s core contribution and the actual differentiator between slop and applied AI.

PAPERS / EVIDENCE

  • Parasuraman & Manzey, Human Factors 52(3), 2010, pp. 381-410. Review of automation complacency + automation bias. Omission errors (miss what automation did not flag) and commission errors (act on automated recommendation against contradictory evidence). Attention is central. Found in both naive and expert participants; cannot be overcome with simple practice; strongest under multiple-task load.
  • Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks & Wilson, CHI 2025. 319 knowledge workers using GenAI at least weekly; 936 first-hand task examples. Higher confidence in GenAI -> less critical thinking; higher self-confidence -> more. Work shifts toward information verification, response integration, task stewardship.
  • Green, Computer Law & Security Review 45 (2022) 105681. VERIFIED FROM THE PAPER ITSELF: 41 policies (a search snippet saying 40 is wrong). Two flaws: (1) evidence suggests people are unable to perform the desired oversight functions; (2) as a result, oversight policies legitimize faulty algorithms, provide a false sense of security, and enable vendors and agencies to shirk accountability.
  • Verga et al., arXiv:2404.18796, 2024. Panel of LLM Evaluators (PoLL): several smaller models from disjoint model families. Across 3 judge settings and 6 datasets, the panel outperforms a single large judge, shows less intra-model bias, and costs over 7x less.
  • Panickssery, Bowman & Feng, NeurIPS 2024, arXiv:2404.13076. GPT-4 and Llama 2 show non-trivial accuracy distinguishing their own output from other LLMs’ and humans’. Linear correlation between self-recognition capability and strength of self-preference. Controlled for straightforward confounders, suggesting self-recognition mechanistically contributes to the bias.
  • Dell’Acqua et al. 2023 (carried from Ch 2): outside the frontier, AI users were 19% less likely to be correct than non-users, via unwarranted trust in confident wrong output.

CONTROLS / LIMITATIONS

Parasuraman & Manzey predates LLMs; it is aviation/process-control automation, applied here by analogy that the chapter should keep explicit. Lee et al. is self-reported and correlational โ€” it cannot establish that AI confidence causes reduced critical thinking. Green is about government algorithms and public-sector policy; the commercial read-across is argued, not measured. Verga and Panickssery are about model judges on benchmark tasks, not about human review quality. None of this measures the seeded-defect practice the chapter recommends; that is an argued recommendation borrowed from aviation and radiology QA, not a result.

THE SHALLOW-PATH GUIDELINES (keep; the chapter offers these without sarcasm)

  1. Diverse panel, not a single judge
  2. Never let the generator grade itself
  3. Randomize position, normalize length
  4. Mechanical checks over “is this good”
  5. Deep random sample (~5%) over universal skim โ€” the sample yields an error-rate estimate
  6. Record the shallow review anyway Limit: lowers error rate, does NOT transfer accountability.

THE ENGINEERED-REVIEW MECHANISMS (each built later)

  1. Reviewable artifacts, not conclusions -> Ch 12
  2. Deterministic checks ahead of attention -> Ch 3, 16
  3. Route attention by irreversibility/novelty/risk -> Ch 15
  4. Preserve raw output to bound the blast radius -> Ch 11
  5. Review record as an artifact with evidence -> Ch 10, 12
  6. Seed known defects and measure your catch rate -> the practice nobody does

DEPENDENCIES

Chapter 2 (intent and verification as jobs that cannot leave the human; the frontier-trust result), Chapter 3 (mechanical checks beat judgments), Chapter 4 (where the verifier bottoms out in a person).

FORWARD BRIDGE

Part 1 ends. Now build the smallest honest version of the stochastic step โ€” a call that names its input, returns its output, and makes failure impossible to mistake for content.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part 1 โ€” Where You Stand

Ten weeks

Week one, the model drafts a paragraph review and you read every line of it. You disagree with two of its four flags, check a source yourself, rewrite one sentence. The whole thing takes twenty minutes and you feel you have earned the output.

Week three, the reviews have been good. You read them properly but you stop re-checking the sources it says it checked. Nothing bad happens.

Week six, you skim. You read the first flag carefully and glance at the rest. You have a feel for what a wrong one looks like now โ€” or you believe you do. Nothing bad happens.

Week ten, you approve.

Not because you got lazy. Because forty consecutive reviews were fine, and on the evidence available to you, reading the forty-first carefully was a poor use of twenty minutes. Every individual decision to look less closely was locally rational. You were updating correctly on the data you had.

This is the trap, and it needs stating plainly because it inverts the usual moral framing:

You are not turned into a rubber stamp by the system’s failures. You are turned into one by its successes.

A system that is wrong 30% of the time keeps you sharp forever; you cannot afford to stop looking. A system that is wrong 2% of the time removes your reason to look โ€” and then the 2% passes through unexamined, signed by you, indistinguishable in the record from work you actually checked.

Call the resulting role what it is. You are a meat proxy: a biological component whose function is to convert machine output into approved output, adding a signature and no information.

What has to be true of a process for a human review to still mean something after the hundredth time?

It has a name, and thirty years of evidence

This is not a new phenomenon and it is not specific to language models. Human factors research has studied it since long before any of this, mostly in aviation and process control, under two names.

Automation bias is the tendency to over-rely on automated recommendations. It produces two distinct error types: omission errors, where you fail to notice a problem because the automation did not flag it, and commission errors, where you act on an automated recommendation in spite of contradictory evidence available to you.

Automation complacency is the attentional counterpart: reduced monitoring of an automated system that has been performing well.

Parasuraman and Manzey’s review of this literature reaches three conclusions that ought to change how you plan your week (Parasuraman & Manzey, 2010):

  1. Complacency and automation bias are manifestations of overlapping, automation-induced phenomena in which attention plays the central role.
  2. It appears in both naive and expert participants in the automation studies Parasuraman and Manzey reviewed โ€” expertise does not protect against complacency in the monitoring task. (That is distinct from the differentiated-learning finding below: experts may still extract more durable skill from AI-assisted practice.)
  3. It cannot be overcome with simple practice. Knowing about it, and resolving to do better, does not fix it.

That third point is the one that should be uncomfortable, because it forecloses the response everyone reaches for first. You cannot solve this with resolve. Nobody in this literature has ever solved it with resolve. If your plan for maintaining review quality is I will be diligent, you do not have a plan.

It also gets worse under exactly the conditions AI creates. Complacency is strongest under multiple-task load โ€” when manual work competes with the monitored task for attention. Which is precisely what happens when a tool makes you nominally responsible for five times as much output.

Measured in this setting

The general finding holds specifically here, too.

Lee and colleagues surveyed 319 knowledge workers who used generative AI at work at least weekly, collecting 936 first-hand examples of real tasks. Two associations came out of it, pointing opposite ways: higher confidence in the AI was associated with less critical thinking, while higher self-confidence was associated with more (Lee et al., 2025).

Note what that pair implies. Trust in the tool and trust in yourself are not the same variable, and they pull against each other. The dangerous configuration is high confidence in the model combined with low confidence in your own judgment โ€” which is also, unfortunately, the configuration that a long run of good outputs manufactures.

The same study found the nature of the work shifting: away from producing, and toward verification and stewardship of information. That is a finding about knowledge work generally, and it is a striking echo of Chapter 2’s four jobs. The work is already moving to where this book says the value is. The question is whether anyone is doing it properly once they get there.

And recall the measured cost of getting it wrong, from Chapter 2: on a task outside current AI capability, consultants using AI were 19% less likely to produce a correct solution than those with no AI at all, because they extended unwarranted trust to confidently-presented wrong output (Dell’Acqua et al., 2023). Not “no better.” Worse. A meat proxy is not a neutral component. It is a component that launders errors into approved results.

Why the model cannot supply the missing part

Here is the structural reason this cannot be delegated back to the system that caused it.

A model optimizes for output that is plausible given its training distribution and whatever preference signal it was tuned on. It has no access whatsoever to your definition of good โ€” not because it is insufficiently capable, but because that definition exists nowhere in its inputs unless you put it there. Chapter 2 listed intent as one of the four jobs that cannot leave the human. This is intent’s failure mode. If nobody supplies “good,” nothing does, and output regresses toward generic competence.

Which gives a definition of slop worth carrying:

Slop is not bad output. Slop is output that is locally plausible and globally unaccountable โ€” nothing in it is wrong enough to catch, and nothing in it was chosen.

That is why slop is so hard to object to in a review meeting. Every sentence is defensible. No sentence is load-bearing. There is no error to point at, because error requires a commitment, and nothing committed to anything. It reads like the average of everything ever written on the topic, because that is approximately what it is.

The corollary is that you cannot detect slop by looking for mistakes. You detect it by asking what the piece claims and what would have to be true for the claim to be wrong. If there is no answer, you are holding slop regardless of how clean the prose is.

Why the failure is catastrophic rather than gradual

The user-facing danger of the meat proxy is not that a bad output ships. Bad outputs ship all the time and get caught.

There are two compounding mechanisms, and the second is the one that hurts.

The first is deskilling โ€” but it is differentiated, not uniform. If you stop doing the work, you can lose the ability to evaluate the work. The review that was cursory because it was unnecessary can become cursory because it is no longer possible. Chapter 2 warned that the path to senior judgment historically ran through doing the junior work. This is the same mechanism, operating on someone who already has the judgment and may quietly spend it down.

But a 2026 field experiment should stop us stating that as a law. Autor and colleagues ran a pre-registered three-month randomized trial with 133 practicing patent lawyers at eleven US firms, using a custom AI drafting assistant, with all work scored by blinded expert attorneys. AI access raised drafting quality at 10 days (+0.34 SD) and 90 days (+0.38 SD), with larger immediate gains for juniors. After three months, everyone redlined an application without AI: treated lawyers outperformed controls by +0.32 SD, but the durable gain was concentrated entirely among seniors (+0.45 SD). Juniors showed no average gain โ€” their scores bifurcated instead, with fewer mediocre outcomes offset by more poor and more good ones (Autor et al., NBER w35720, 2026).

Read carefully, with the authors’ limits: one occupation, Google-funded, a working paper rather than peer-reviewed publication. Measured: AI-assisted practice built lasting judgment for seniors (+0.45 SD on the unaided redline) while juniors bifurcated with no average gain. The authors’ reading: foundational expertise may be a prerequisite for extracting durable skill from AI-assisted practice. The book’s extension, not theirs: if so, the one-way door is realest where expertise was thinnest to begin with. That refines rather than removes the warning.

Field observation of developers points the same way. Experienced developers using coding agents retain architecture and design control, supply context, and repeatedly verify rather than delegate โ€” control as the observed expert pattern, from field observations (N=13) and surveys (N=99) (Huang et al., 2025). And across roughly 400,000 Claude Code sessions, Anthropic finds people disproportionately make planning decisions (what, ~70%) while the model makes execution decisions (how, ~80%), with verified success rising from ~15% for novice-rated sessions to ~28โ€“33% for intermediate and above โ€” vendor observational data, classifier-judged, not causal, but consistent: the human owns intent, and expertise amplifies what the agent can do (Anthropic, 2026).

The second is the missing trail. When the failure is finally detected โ€” externally, by a customer, an auditor, a regulator, or a court โ€” the question is never “is this one wrong?” The question is how far back does this go?

And you cannot answer it. Your record is a sequence of approvals. The log shows a timestamp and a signature on each one. Open any entry and you find the edited file, but not the bytes the model returned, not the checklist of what was verified, only the source link pointing at whatever version is current now. An auditor sampling across ten thousand such entries finds the same thin record repeated, and has to treat the ones you read and the ones you glanced at as the same thing.

That is the catastrophic failure the meat proxy is walking toward. An unbounded blast radius, because the process that produced ten thousand results retained no evidence about any of them.

This is why the runtime chapters of this book are not administrative overhead. A record of what was checked and accepted (Chapters 14 and 16), artifacts that preserve what was actually returned (Chapters 11 and 17), and claims carrying their evidence (Chapter 18) are precisely the machinery that converts “everything is suspect” into “these four hundred are suspect, and here is why.”

The loop that produces the proxy, and where it exits into catastrophe:

    flowchart TD
    S["repeated success<br/><i>forty good outputs in a row</i>"] --> T["trust rises<br/><i>each step locally rational</i>"]
    T --> C["checking thins<br/><i>skim, then approve</i>"]
    C --> U["failure passes unexamined<br/><i>signed by you</i>"]
    U -.->|"undetected failure<br/>looks like more success"| S
    U --> E["approvals lose evidential value<br/><i>no record of what was checked</i>"]
    E --> B["unbounded blast radius<br/><i>how far back does this go?</i>"]
  

The institutional trap

There is a reason the meat proxy is not merely tolerated but frequently rewarded, and it is worth being clear-eyed about.

Ben Green surveyed 41 policies requiring oversight of government algorithms and identified two flaws. The first is that the evidence indicates people are unable to perform the oversight functions the policies assume they can perform. The second is worse: human oversight requirements legitimize the use of faulty and controversial algorithms without addressing their underlying problems, providing a false sense of security and allowing vendors and agencies to shirk accountability (Green, 2022).

Read that in a commercial setting. When a policy requires a human in the loop, the organization needs a human signature. It does not need, and often cannot afford, genuine human review. So the role is created, the box is ticked, and the human’s actual function is to be the entity that absorbs responsibility when something goes wrong. The meat proxy is not a failure of the oversight regime. In a great many deployments, the meat proxy is the oversight regime.

This is the honest answer to the observation that some people will simply be meat proxies and nothing will stop them. That is true, and it is not primarily a character flaw. It is an equilibrium: the role pays, the work is easy, the failure is deferred and diffuse, and the institution frequently prefers the signature to the scrutiny. Expect it to persist, and expect it to be comfortable for a long time.

If you are going to do it anyway

Some readers will do this regardless, and some tasks genuinely do not justify a careful human pass. Pretending otherwise helps nobody. So: if your review is going to be shallow, here is how to make a shallow review as strong as it can be. This section is offered without sarcasm. A well-built shallow review beats a badly-built one by a large margin.

Use a panel, not a judge. Verga and colleagues evaluated a Panel of LLM Evaluators โ€” several smaller models drawn from disjoint model families โ€” against a single large judge, across three judge settings and six datasets. The panel outperformed the single large judge, exhibited less intra-model bias, and cost over seven times less (Verga et al., 2024). The diversity is doing the work here, not the size.

Never let the generator grade its own output. LLM judges favor their own generations, and Panickssery and colleagues found a linear relationship between a model’s ability to recognize its own output and the strength of its preference for it (Panickssery et al., 2024). Self-evaluation is not a weak check; it is a check whose bias points in exactly the wrong direction.

Control the known artifacts. Judge models exhibit position bias and verbosity bias. Randomize the order of candidates and normalize for length, or you are measuring presentation rather than quality.

Prefer mechanical checks to judgments. This is Chapter 3’s rule applied to review. Do not ask a model “is this good.” Ask deterministic questions: does every cited source resolve? Does every quoted sentence appear verbatim in the source? Do the numbers sum? Does it compile? Are all flagged spans real offsets into the document? Each of these is free, exact, and repeatable, and each one removes a class of failure from the pile you are eyeballing. Exact is not the same as sufficient: a check is precise about what it tests and silent about everything else, and Chapter 21 shows a well-built one accepting a grounded, verbatim-quoted, wrong answer.

Sample deeply instead of skimming universally. A genuine, unhurried review of a random 5% is worth more than a glance at 100%, for a reason that is not obvious: the random sample produces an error rate estimate. Skimming produces nothing you can act on. One of these tells you when the system degrades; the other tells you that you looked.

Record the shallow review anyway. Version, checks run, and approver. This costs seconds and it is the difference between a bounded incident and an unbounded one.

And the honest limit on all of it: this lowers your error rate. It does not transfer accountability. If the output matters โ€” if someone is relying on it, if it is going in front of a regulator, if it is load-bearing for a decision โ€” none of these substitutes for a person who understands the work. The panel is a filter, not a signatory.

The real path: review is architecture, not willpower

Now the version for work you actually care about.

The research says diligence does not survive a system that is usually right, in experts, with practice. So stop treating review as a virtue and start treating it as a design problem. The question is not will I be careful? It is: what would have to be true of this process for a careful review to be cheap, structurally unavoidable, and measurably effective?

Six mechanisms, each of which the rest of this book builds:

Make the model emit reviewable artifacts, not conclusions. A claim with a source span and a quoted sentence can be verified in four seconds. Three paragraphs of confident prose cannot be verified at all โ€” only agreed with. This single change does more for review quality than any amount of attention, because it converts an act of judgment into an act of checking (Chapter 18).

Put deterministic checks in front of your attention, not beside it. Your attention is the scarcest and most expensive component in the system. Spend it only on what a machine could not have decided. Everything the machine could decide should already be decided by the time you look (Chapters 3, 14 and 21).

Route attention by risk, not by volume. Do not review evenly. Review what is irreversible, novel, or high-consequence. The authority machinery in Chapter 20 exists precisely to identify these, and its real purpose is as much about directing human attention as about preventing action.

Preserve the raw output. Not the edited version โ€” the bytes the model actually returned, content-addressed, before anyone touched them (Chapters 11 and 17). This is what converts a discovered failure into a bounded one.

Make the review record an artifact with evidence attached. An approval with no evidence is not a review, it is a signature. A review record should say what was checked and what the check returned (Chapters 14 and 18).

Measure your own review. This is the one almost nobody does, and it is the only way to find out whether you have already become a proxy. Periodically seed known defects into the queue โ€” a fabricated citation, a flag pointing at a sentence that is not there, a number that does not add up โ€” and record whether you caught them. Aviation and radiology both do versions of this. Your catch rate on seeded defects is the only honest measurement of whether your review still means anything, and it will be lower than you expect the first time you run it. The later chapters of this book hold their automated verifiers to the same standard: each is run against deliberately corrupted copies of its own evidence, and a verifier that misses a seeded corruption has failed.

Notice what these have in common. Not one of them asks you to try harder. Each one changes the shape of the work so that the careful path is the cheap path. That is the only intervention the evidence supports.

The heart at the end of the machine

The model generates. The runtime coordinates. The tools act. The verifiers check what can be checked mechanically.

What is left โ€” the definition of good, and the enforcement of it against a system that will happily produce plausible mediocrity forever โ€” is yours. It is not a formality at the end of the pipeline. It is the only part of the pipeline that determines whether any of the rest of it produced something worth having.

That is the whole difference between slop and applied AI. Not the model. Not the prompt. Not the framework. The heart at the end of the machine, and whether it is still beating by week ten.

Do this now

Thirty minutes, and it will be uncomfortable. Measure your own catch rate.

  1. Take ten recent AI outputs you approved. Real ones.
  2. Have someone else โ€” or a script โ€” corrupt three of them: change a cited year, swap a quoted sentence for one that does not appear in the source, alter a number so a total no longer sums.
  3. Review all ten the way you normally would. At your normal speed. Not the way you would review them knowing this was a test โ€” that is the whole point, and it is why someone else should introduce the defects.
  4. Record how many of the three you caught.

Then write the number down with the date. It is the only honest measurement of whether your review still means anything, and it is the one number in this book that nobody else can produce for you.

Failure modes

  • Assuming diligence is a plan. Complacency appears in experts and does not yield to practice.
  • Reading success as safety. The better the system performs, the faster you stop checking, and the more your attention is needed for exactly the cases you are no longer seeing.
  • Letting the generator judge itself. Self-preference scales with self-recognition; the bias points the wrong way.
  • A single judge model. A panel of diverse smaller models beat it, with less intra-model bias, at a seventh of the cost.
  • Asking “is this good” instead of “does this check out.” Judgment where a mechanical check was available.
  • Skimming everything instead of sampling deeply. Produces no error estimate, so degradation is invisible until it is external.
  • Approving without evidence. Converts a recoverable incident into an unbounded audit.
  • Confusing compliance with oversight. A required signature is not a performed review, and the policy may exist to place blame rather than to catch errors.

What this chapter established

  • The meat proxy is produced by a system’s successes, not its failures. Every step down the slope is locally rational.
  • It is a documented phenomenon with thirty years of evidence: automation bias (omission and commission errors) and automation complacency, present in experts, not overcome by practice, worse under multi-task load.
  • Measured here: higher confidence in AI is associated with less critical thinking, while higher self-confidence is associated with more; and unwarranted trust near the capability frontier made outcomes worse than no AI.
  • The model cannot supply the missing part, because your definition of good is not in its inputs. Slop is output that is locally plausible and globally unaccountable.
  • The catastrophic failure mode is not a bad result โ€” it is an unbounded blast radius, caused by a process that approved ten thousand results and retained evidence about none.
  • Deskilling is differentiated (measured: seniors gained durable unaided skill, juniors bifurcated; the authors read this as expertise being a prerequisite). The book’s extension: the one-way door may therefore be realest where expertise was thinnest. Neither establishes universal upskilling.
  • Meat proxies are institutionally rewarded: oversight requirements can legitimize weak systems and shift accountability without improving decisions.
  • If reviewing shallowly: diverse panel, never self-evaluation, control position and length, mechanical checks over judgments, deep random sample over universal skim, and record it. This lowers error rates; it does not confer accountability.
  • If reviewing seriously: reviewable artifacts, deterministic checks ahead of attention, attention routed by irreversibility, preserved raw output, review records carrying evidence, and seeded defects to measure your own catch rate.
  • Review is architecture, not willpower. It is also your core contribution.

Next

One question remains before we build, and it is the one that decides how much of any of this you can afford to do.

Every mechanism in this chapter costs something. Deeper review costs attention. Panels cost tokens. Preserved artifacts cost storage. And the thing being reviewed costs money on every single call, forever โ€” which is a strange property for software to have, because software used to get finished. The next chapter works out what you are actually buying when you buy intelligence, why a verifier puts a ceiling on that price and its absence removes one, and how to take the cost back out.

Continue with The Price of Intelligence.

References

  • Raja Parasuraman and Dietrich H. Manzey. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, vol. 52, no. 3 (2010), pp. 381โ€“410. https://doi.org/10.1177/0018720810376055
  • Hao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ‘25). https://doi.org/10.1145/3706598.3713778
  • Ben Green. The Flaws of Policies Requiring Human Oversight of Government Algorithms. Computer Law & Security Review, vol. 45 (2022), 105681. https://doi.org/10.1016/j.clsr.2022.105681
  • Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv:2404.18796, 2024. https://arxiv.org/abs/2404.18796
  • Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM Evaluators Recognize and Favor Their Own Generations. Advances in Neural Information Processing Systems 37 (NeurIPS), 2024. https://arxiv.org/abs/2404.13076
  • Fabrizio Dell’Acqua et al. Navigating the Jagged Technological Frontier. Harvard Business School Working Paper 24-013, 2023. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
  • David Autor, Tanya Rodchenko, Josh Martin, Zanna Iscenko, Scott Strand, David Pearl, and Melissa Ferere. Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting. NBER Working Paper 35720, September 2026. https://www.nber.org/papers/w35720 โ€” Pre-registered RCT, 133 patent lawyers; +0.34/0.38 SD assisted quality, +0.32 SD unaided redline concentrated in seniors (+0.45 SD), juniors bifurcated. Working paper, single occupation, Google-funded.
  • Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, and Brian Hempel. Professional Software Developers Don’t Vibe, They Control. arXiv:2512.14012, 2025 (v2 Aug 2026). https://arxiv.org/abs/2512.14012 โ€” Field observations N=13, survey N=99; experts retain design control, supply context, verify.
  • Anthropic. Agentic coding and persistent returns to expertise. June 2026. https://www.anthropic.com/research/claude-code-expertise โ€” ~400,000 Claude Code sessions; user owns ~70% planning, model ~80% execution; verified success 15% novice vs 28โ€“33% intermediate+. Vendor observational data, classifier-judged success.