Never Stand in Front of the Steamroller
Part 1 β Where You Stand
A dated position
This book will be wrong about things.
Not as a disclaimer. As a statement about what kind of document this is. Everything here is a view from one position on a hill that is still being climbed. From where we are standing in late 2026, certain shapes are visible and certain ones are not, and some of what currently looks like a permanent feature of the landscape will turn out to have been a cloud.
The useful question is not whether a book about AI will age badly. It is whether it ages badly in a way you can detect and correct, or in a way that quietly misleads you for two years.
So rather than open with a hedge, here is what being wrong responsibly actually looks like, from a lab that does this well.
In July 2025, METR published a randomized controlled trial on AI coding tools. Sixteen experienced open-source developers, 246 real tasks in their own repositories β mature projects averaging over a million lines of code β each task randomly assigned to allow or disallow AI assistance. The developers forecast beforehand that AI would cut their completion time by 24%. Afterwards, having done the work, they estimated it had cut their time by 20%. Measured, allowing AI increased completion time by 19% (METR, 2025).
That result traveled a long way, and deservedly: a 39-point gap between what skilled practitioners believed about their own productivity and what was measured is a serious finding.
Then, in February 2026, METR published an update saying their experiment design had a problem. Developers were increasingly declining to participate in the no-AI condition, even at $50/hour, and between 30% and 50% of participants were avoiding submitting the very tasks they expected AI to accelerate. That is selection bias pointing directly against the headline. Their newer estimates moved the other way β around β18% time for returning developers (confidence interval β38% to +9%) and β4% for newly recruited ones (β15% to +9%) β with both intervals crossing zero. METR’s own summary is that developers are likely more sped up in early 2026 than their early-2025 estimate suggested, and that the new data is only very weak evidence (METR, 2026).
Read what happened there. A careful lab ran a clean experiment, got a striking number, published it, kept looking, found the flaw themselves, and stated plainly how weak the replacement evidence is. Nothing was retracted and nothing was defended past its evidence.
That is the standard this book holds itself to, and it is why Part 5 spends five chapters on experiments, including one in which a result the author wanted to be true did not survive replication. The durable lesson from METR is not the 19%. It is the 39-point gap between belief and measurement β which survived the design correction, and which is exactly why this book will not let a model, or a person, grade its own homework.
Given that the ground is moving, what can you actually stand on?
The one thing that looks clear
Here is the claim this chapter exists to make.
Never stand in front of the steamroller.
The steamroller is not “AI.” That is too vague to act on. The steamroller is a specific, repeating process:
flowchart LR
T["tacit<br/><i>known by doing</i>"] --> C["codified<br/><i>written down, teachable</i>"]
C --> A["automated<br/><i>executed by machine</i>"]
A --> M["commodity<br/><i>priced near zero</i>"]
style A fill:#00000000,stroke-dasharray: 4 4
Knowledge starts tacit. Over time it gets codified β written into documentation, patterns, Stack Overflow answers, training corpora. Once codified, it becomes automatable. Once automated, it is priced accordingly.
This process is old. What changed is the speed of the last two steps, and the fact that a model trained on the codified layer can execute it immediately across every domain at once, without anyone having to build a domain-specific tool.
To stand in front of the steamroller is to hold a role whose value is the codified part. If what you are paid for is knowing things that are written down and executing procedures that are described, you are standing on the section of road the machine is currently flattening.
This is not a prediction that programmers will be unemployed. It is narrower and better supported than that, and the narrow version is the one worth acting on.
What the payroll data says
Brynjolfsson, Chandar, and Chen have been tracking this in administrative payroll microdata from ADP covering millions of US workers. Their August 2026 revision, using data through June 2026, reports six facts. The relevant ones:
- There is no evidence of widespread, economy-wide job displacement. That is their first finding, and it should be the first thing anyone quoting this work says.
- Employment of workers aged 22β25 in the most AI-exposed occupations now stands about 19% below where it would be had it kept pace with similarly aged workers in less-exposed occupations. In levels: employment for that age group in the two most exposed quintiles fell about 11% between November 2022 and June 2026, while the same age group in the three least-exposed quintiles grew about 10%.
- Experienced workers show no comparable gap.
- The divergence has widened steadily β 15% at the July 2025 data vintage, 19% by June 2026.
- It operates primarily through reduced hiring rather than increased separations.
- Declines concentrate in occupations where AI usage substitutes for human tasks. Where usage complements workers, employment is flat or rising, especially for experienced workers.
And the mechanism finding, which is the one this whole chapter turns on:
Employment declines for young workers appear in occupations that involve codified knowledge. Occupations that involve tacit knowledge see faster employment growth for experienced workers (Brynjolfsson, Chandar & Chen, 2026).
That is the steamroller, measured, in the authors’ own vocabulary rather than mine.
Now bound it, because the authors do. They describe these as early, descriptive indicators β canaries in the coal mine β not causal estimates. The patterns attenuate when controlling for education. Some divergent trends predate generative AI, particularly around the pandemic. The effects are more pronounced in the ADP analysis sample than in national survey benchmarks. And again: no economy-wide displacement.
Note also what “reduced hiring rather than separations” means concretely. The steamroller is not running people over. It is declining to lay track where the next cohort was going to walk. That is a materially different problem, and I will come back to why it is the hardest part of this chapter’s advice.
Why “be excellent at the codified part” stopped working
The instinctive response to all this is to get better. Be a stronger engineer than the model. That response has been measured too, and it fails for a reason that is not obvious.
Brynjolfsson, Li, and Raymond studied the staggered rollout of a generative AI assistant across 5,179 customer support agents. Access raised productivity β issues resolved per hour β by 14% on average. But the average hides the finding: a 34% improvement for novice and low-skilled workers, and minimal impact on experienced, highly skilled workers. Their suggestive mechanism is that the model disseminates the best practices of the most able workers, moving newer workers down the experience curve faster (Brynjolfsson, Li & Raymond, 2025).
Read that from the perspective of the experienced worker. The tool took the codified portion of your expertise, the part that could be extracted from transcripts of your best work, and distributed it to everyone who did not have it. You gained little. The gap between you and a novice narrowed sharply.
That is not replacement. It is compression. If your market value came from being better than average at the part of the job that can be written down, the model did not take your job β it took your differential.
This is why the entry-level effect and the compression effect are the same story seen from two ends. The codified layer is being commoditized. Juniors are affected first because their roles are the most purely codified. Seniors are affected second and less visibly, through the erosion of whatever part of their premium was codifiable.
Bound this one too: one firm, one occupation, a support context with unusually clean productivity metrics and unusually codifiable expertise. Software engineering is not customer support. The direction is what transfers, not the 34%.
A newer result complicates the compression story without removing it. In a pre-registered three-month trial with 133 patent lawyers, AI assistance raised drafting quality (+0.34 SD at 10 days, +0.38 SD at 90 days, larger for juniors) β but durable unaided judgment afterwards concentrated in seniors (+0.45 SD), while juniors bifurcated rather than improving on average (Autor et al., NBER w35720, 2026). That is NBER working-paper evidence, one occupation, Google-funded: it supports, within those conditions, that foundational expertise may be a prerequisite for extracting lasting skill from AI-assisted practice. Compression of immediate output and differentiated learning can both be true β and the ladder problem below gets harder, not easier, if juniors get the output gain without the judgment gain.
Where the ground is solid: the jagged edge
If the codified part is being flattened, what is not?
Dell’Acqua and colleagues ran a preregistered field experiment with 758 BCG consultants, roughly 7% of the firm’s individual-contributor consultants. On 18 realistic tasks chosen to sit inside current AI capability, consultants using AI completed 12.2% more tasks, 25.1% faster, at significantly higher quality. On a complex task deliberately chosen to sit outside that capability, consultants using AI were 19% less likely to reach a correct solution than those without it β because they extended unwarranted trust to confidently-presented, substantively wrong output (Dell’Acqua et al., 2023).
The authors named the shape of the problem: a jagged technological frontier. Some tasks fall easily inside current capability. Others, apparently similar in difficulty, fall outside it. The boundary is irregular and it is not visible from the task description.
So here is the job that has value: knowing where the edge is. And notice its properties. It cannot be codified, because the frontier moves β which means it cannot be handed to the model, and it cannot be learned once and banked. It is a standing obligation to check. The same paper shows what happens to people who skip it: they do worse than people with no AI at all.
That is the first of four positions worth holding, and it is the one this book’s Chapter 3 turns into an engineering rule rather than an intuition.
The four jobs the model cannot take
Chapter 1 ended with an assignment of responsibilities: model = cognition, runtime = coordination, tools = action, verifiers = evidence, human = intent and authority. That was an architecture diagram. Read it again as a map of where to stand.
| The job | What it is | Why the model cannot hold it |
|---|---|---|
| Intent | Deciding what should be done, and what would count as done | The model can propose objectives; it cannot be the thing that wants one. Someone must own the definition of success and be accountable for it. |
| Authority | Deciding who is permitted to cause which effects | Capability is not authority (Chapter 20). A system that authorizes its own actions has no authorization step, only a delay before one. |
| Verification | Obtaining evidence about reality, not agreement from another model | Verification requires contact with the world β a test run, a source checked, a result reproduced (Chapter 21). Consensus among generators is not evidence. |
| Frontier judgment | Knowing when the stochastic component is the right tool at all | The frontier is jagged and moving; the model is measurably confident on the wrong side of it (Dell’Acqua et al., 2023). |
The property these four share is the important one. They do not get cheaper as models improve. They get more valuable β because a better model produces more output per hour, and every unit of that output needs intent behind it, authority over it, and verification after it. Capability growth increases the volume of work requiring governance while doing nothing to supply the governance.
That is what it means to not stand in front of the steamroller. Not “move into management,” which is advice about a job title. It means occupying the part of the process whose demand is created by the thing doing the flattening.
And this is the point at which the career argument and the architecture argument turn out to be one argument. Every remaining chapter of this book builds machinery for exactly these four jobs: explicit context and budgets so intent is legible (Chapters 14β15), a ledger and artifact store so what happened survives (Chapters 16β17), claims and evidence levels so support stays distinguishable from assertion (Chapter 18), capability kept separate from authority (Chapter 20), independent verification (Chapter 21), and a router that holds the judgment about when to spend a call at all (Chapter 28).
Applied AI is not a set of tricks for getting more out of a chat window. It is the engineering discipline of the position you want to be standing in.
The same four jobs decide something closer to home. Once an assistant can build software, deciding what your own tools should do, what they may change, and how you would know a change helped is intent, authority and verification applied to your working environment. Chapter 30 returns to that.
Where this argument is weakest
A chapter that opens by promising to be wrong owes you its own best counter-arguments. Here are the ones I find genuinely hard.
The ladder problem is real and this chapter does not solve it. The route to senior judgment has historically run through the junior codified work, where you learned where the frontier was by being wrong about it on small things for several years. If the steamroller removes the bottom rungs β and “reduced hiring rather than separations” says precisely that it does β then “acquire senior judgment” is not actionable advice for someone who cannot get hired to acquire it. I do not have a good answer, because this is a structural problem requiring a structural response, and individuals absorbing it as a personal failure are misreading it.
The composition problem. Advice that works for one person can be arithmetically impossible for a cohort, since if everyone moves to directing AI, the ratio of directors to directed work becomes absurd. “Get out of the way of the steamroller” describes a smaller destination than the road it is flattening, and anyone telling you otherwise is selling something.
Verification might automate further than I am betting. This book’s central wager is that verification resists automation because it requires contact with reality, which is a bet rather than a theorem. Automated test generation, formal methods, and simulation all push on it, and if verification collapses into the model, a good part of this chapter’s advice collapses with it.
The evidence is young and mostly descriptive. The strongest labor result here is explicitly labeled by its authors as descriptive rather than causal, attenuates under education controls, and runs larger in its sample than in national benchmarks, while the productivity results come from single firms or single occupations. The METR reversal in this chapter’s own opening is a live demonstration that confident readings of this literature can age badly within months.
The metaphor invites fatalism. “Steamroller” can be read as “nothing you do matters.” That is the opposite of the point. The point is that where you stand is a decision, it is available to you, and the evidence about which ground is solid is better than it was two years ago.
How to read the rest of this book
Claims in this book come in different strengths, and the book tries never to let a weaker one borrow the authority of a stronger one:
- Measured β a number produced by an experiment, reported with its sample and its bounds.
- Inspected β an observation about source code that was actually read, identified by file and symbol.
- Argued β a position defended from stated premises, like the determinism rule in Chapter 3.
- Predicted β a claim about how things will go. The steamroller is one of these. It is the weakest category, and everything in it should be held loosely.
In the early chapters the strength is stated in the prose. From Chapter 19 onward, where the book leans on its own experiments, claims about CodeAI also carry explicit tags that separate a pinned measurement from a demonstration that ran but was not pinned, a frozen historical report, code that was only read, and evidence that is still owed.
Where a number appears, it has a source and a stated limit. Where this book proposes a contract stronger than the code currently implements, it says so. Where an experiment failed, the failure is the chapter.
Do this now
Fifteen minutes, one page. Audit your own exposure.
- List the five things you were paid to do last month. Be concrete β tasks, not a job title.
- For each, mark whether its value is the codified part (knowing what is written down, executing a described procedure) or the tacit part (judgment about what should be done, or about whether it worked).
- For each codified one, write the sentence: if a model did this at 90% quality for a hundredth of the cost, what would still need me? If you cannot finish the sentence, that is the finding.
- Now mark which of this chapter’s four jobs β intent, authority, verification, frontier judgment β you actually hold, as opposed to being adjacent to.
To turn the audit into a number, score it as a constructed teaching sketch (illustrative tasks, not measured data). Mark each task C (value is the codified part) or T (value is tacit judgment), then count codified tasks where you hold none of the four jobs:
| Example task | C / T | Four-job held? | Exposed? |
|---|---|---|---|
| Reset passwords per runbook | C | none | yes |
| Triage routine support tickets | C | none | yes |
| Write weekly status summary | C | intent | no |
| Judge an ambiguous outage escalation | T | frontier judgment | no |
| Sign off a production migration | T | authority, verification | no |
Here 2 of 5 tasks are exposed (2 codified tasks with no job held). The reader can now do one new thing after this chapter: compute their own exposed count and name the boundary β any task scoring C with no job held is the section of road to move off first, before any career decision is made on vibes.
Keep the page. Chapter 7 asks you to measure things; this is the only measurement in the book where you are the instrument and the subject.
Failure modes
- Reading “no economy-wide displacement” as “nothing is happening.” Both facts come from the same paper. The aggregate is calm; the 22β25 cohort in exposed occupations is not.
- Reading the entry-level data as a reason to disparage juniors. The mechanism is reduced hiring in codified roles, not a judgment about the people who would have filled them.
- Trusting confident output near the frontier. Measurably worse than not using the tool at all on the wrong side of the edge.
- Believing your own productivity estimate. Skilled developers were off by 39 points about themselves, in the direction of optimism.
- Treating “move up the stack” as a completed action. Frontier judgment is a standing obligation to re-check, because the frontier moves.
- Quoting any number in this chapter without its bound. Every one of them is from a specific sample in a specific period.
What this chapter established
- This is a dated position, and being wrong responsibly means publishing the correction as loudly as the finding β the standard METR set for itself between 2025 and 2026.
- The steamroller is the tacit β codified β automated β commodity pipeline, and the danger is holding a role whose value is the codified part.
- Measured: no economy-wide displacement, but a 19% kept-pace shortfall for 22β25-year-olds in AI-exposed occupations, driven by reduced hiring, concentrated where AI substitutes rather than complements, and specifically in occupations involving codified knowledge β all labeled descriptive rather than causal by its authors.
- Getting better at the codified part does not escape it: AI compressed the skill distribution by handing the best workers’ codifiable expertise to novices, with 34% gains at the bottom and minimal gains at the top.
- The frontier is jagged and invisible, and on the wrong side of it AI users did worse than non-users.
- Four jobs the model cannot hold: intent, authority, verification, frontier judgment. They appreciate rather than depreciate as capability grows.
- The honest weaknesses: the broken ladder, the composition problem, the possibility that verification automates, and evidence that is young and descriptive.
Next
Three of those four jobs β intent, authority, verification β are engineering problems, and the rest of this book builds them. The fourth, frontier judgment, needs a rule before it can be built into anything, because “use AI where it’s good” is not something you can put in a file.
The next chapter turns it into one. It argues that a component belongs on the deterministic side unless it demonstrably cannot be, that doubt about which side is itself the answer, and that the stochastic box is more stochastic than its configuration claims β temperature=0 does not make a model deterministic, and the reason has nothing to do with sampling.
Continue with If There’s Any Doubt, It’s Deterministic.
References
- Erik Brynjolfsson, Bharat Chandar, and Ruyu Chen. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab, August 2026 revision (data through June 2026). https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/
- Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond. Generative AI at Work. The Quarterly Journal of Economics, vol. 140, no. 2 (2025), pp. 889β942. https://doi.org/10.1093/qje/qjae044
- Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, FranΓ§ois Candelon, and Karim R. Lakhani. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Harvard Business School Technology & Operations Mgt. Unit Working Paper No. 24-013, 2023. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321
- Joel Becker, Nate Rush, Elizabeth Barnes, and David Rein (METR). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089, July 2025. https://arxiv.org/abs/2507.09089
- METR. We are Changing our Developer Productivity Experiment Design. February 24, 2026. https://metr.org/blog/2026-02-24-uplift-update/
- David Autor, Tanya Rodchenko, Josh Martin, Zanna Iscenko, Scott Strand, David Pearl, and Melissa Ferere. Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting. NBER Working Paper 35720, September 2026. https://www.nber.org/papers/w35720