← Applied AI

Never Stand in Front of the Steamroller

Machines have taken the codified part of skilled work before, and no amount of skill at that part saved the job. What is new is the pace. The roles being flattened are the ones whose value is the codified part; this book states that as a dated position and says how it could be wrong. The four jobs this book keeps outside the model are the ones worth holding.

Part 1 — Where You Stand

The plain version

The historical premise of this chapter is not controversial: machines have displaced codified parts of skilled work before. The inference for AI — which parts are now exposed, how fast the boundary moves, and where durable human responsibility remains — is the claim this chapter has to earn.

“Computer” was a job title before it was a machine: rooms of people did calculations by hand for engineers, observatories and space programs. Typesetters set newspapers line by line, first by hand and later on hot-metal machines, until photocomposition and then desktop publishing took the trade. Keypunch operators turned programs and data into stacks of punched cards. Typing pools produced every letter an office sent. Each of those was skilled work, done by people who were good at it and some who were superb. None of that skill protected the job once a machine could do the codified part.

The best-measured case is the telephone operator. In the early 1900s it was among the most common jobs for American women, and between 1920 and 1940 AT&T replaced operators with mechanical switching across more than half of the US telephone network. Feigenbaum and Gross traced what followed. Incumbent operators were hit hardest: a decade later they were more likely to be in lower-paying occupations or no longer working. Later cohorts of young women were not employed less overall; they went into clerical and service jobs, including new kinds of work (Feigenbaum & Gross, 2024). Keep that pattern in mind. The current data differs from it in a way that matters.

What is new is speed. METR measures the length of software tasks, timed by how long they take skilled people, that frontier models can complete with 50% success. That length has doubled roughly every seven months since 2019, and the authors note the trend may have accelerated in 2024 (Kwa et al., 2025).

Bound it: one task suite, mostly software, a 50% success rate that is a long way from reliable, and a trend rather than a law — it can slow. But take it at face value for a moment. Doubling every seven months is roughly a thousandfold in six years.

Now the plain logic. If what you are good at is the codified part of a job — the part that can be written down, shown by example and checked — you are competing with a capability whose measured frontier has been moving quickly and whose deployment can spread much faster than a new human skill. Once that codified part can be automated cheaply and adequately, excellence at it alone is not durable protection. Being among the best typesetters in 1975 was a real achievement, and it was not a career plan. The useful question is not whether you are better than the model today. It is which parts of your value remain when the codified part becomes cheap.

Never stand in front of the steamroller.

That is the whole chapter in one line. The rest makes it precise enough to act on.

Given that the ground is moving, what can you actually stand on?

What the steamroller is

The steamroller is not “AI.” That is too vague to act on. The steamroller is a specific, repeating process:

    flowchart LR
    T["tacit<br/><i>known by doing</i>"] --> C["codified<br/><i>written down, teachable</i>"]
    C --> A["automated<br/><i>executed by machine</i>"]
    A --> M["commodity<br/><i>priced near zero</i>"]
    style A stroke-dasharray: 4 4
  

Knowledge often starts tacit. Some of it gets codified — written into documentation, patterns, Stack Overflow answers, training corpora. Codification makes automation easier because examples, procedures, and checks become available to machinery. When automation becomes reliable and cheap enough, the market value of executing that codified procedure can fall sharply.

This process is old; the typesetters and telephone operators above rode versions of it. What changed is the speed and breadth with which general-purpose models can attempt codified tasks across many domains without a bespoke model being built for each one.

To stand in front of the steamroller is to hold a role whose value is concentrated in the codified part. If most of what you are paid for is retrieving written knowledge and executing described procedures, that portion of the role is the portion this chapter treats as exposed.

This is not a prediction that programmers will be unemployed, nor a claim that codification guarantees automation. It is narrower: codified, checkable work is the part for which the path to cheap automation is shortest.

What the payroll data says

Brynjolfsson, Chandar, and Chen have been tracking this in administrative payroll microdata from ADP covering millions of US workers. Their August 2026 revision, using data through June 2026, reports six facts. The relevant ones:

  • There is no evidence of widespread, economy-wide job displacement. That is their first finding, and it should be the first thing anyone quoting this work says.
  • Employment of workers aged 22–25 in the most AI-exposed occupations now stands about 19% below where it would be had it kept pace with similarly aged workers in less-exposed occupations. In levels: employment for that age group in the two most exposed quintiles fell about 11% between November 2022 and June 2026, while the same age group in the three least-exposed quintiles grew about 10%.
  • Experienced workers show no comparable gap.
  • The divergence has widened steadily — 15% at the July 2025 data vintage, 19% by June 2026.
  • It operates primarily through reduced hiring rather than increased separations.
  • Declines concentrate in occupations where AI usage substitutes for human tasks. Where usage complements workers, employment is flat or rising, especially for experienced workers.

And the mechanism finding, which is the one this whole chapter turns on:

Employment declines for young workers appear in occupations that involve codified knowledge. Occupations that involve tacit knowledge see faster employment growth for experienced workers (Brynjolfsson, Chandar & Chen, 2026).

That is the steamroller, measured, in the authors’ own vocabulary rather than mine.

Now bound it, because the authors do. They describe these as early, descriptive indicators — canaries in the coal mine — not causal estimates. The patterns attenuate when controlling for education. Some divergent trends predate generative AI, particularly around the pandemic. The effects are more pronounced in the ADP analysis sample than in national survey benchmarks. And again: no economy-wide displacement.

Note also what “reduced hiring rather than separations” means concretely. The steamroller is not running people over. It is declining to lay track where the next cohort was going to walk. That is a materially different problem, and I will come back to why it is the hardest part of this chapter’s advice.

Compare the telephone operators. There, incumbents were hurt and the next cohort found other work. Here, so far, experienced incumbents show no gap, and it is the next cohort that is not being hired. Whether that cohort finds other work, as the young women of the 1930s did, is exactly what the data cannot yet tell us.

Why “be excellent at the codified part” stopped working

The instinctive response is to get better — be a stronger engineer than the model. The opening said that at the codified part it does not matter how good you are. That sounds like rhetoric; the closest measurement shows what it looks like in practice, and the mechanism is not the obvious one.

Brynjolfsson, Li, and Raymond studied the staggered rollout of a generative AI assistant across 5,179 customer support agents. Access raised productivity — issues resolved per hour — by 14% on average. But the average hides the finding: a 34% improvement for novice and low-skilled workers, and minimal impact on experienced, highly skilled workers. Their suggestive mechanism is that the model disseminates the best practices of the most able workers, moving newer workers down the experience curve faster (Brynjolfsson, Li & Raymond, 2025).

Read that from the perspective of the experienced worker. The tool took the codified portion of your expertise, the part that could be extracted from transcripts of your best work, and distributed it to everyone who did not have it. You gained little. The gap between you and a novice narrowed sharply.

That is not replacement. It is compression. If your market value came from being better than average at the part of the job that can be written down, the model did not take your job — it took your differential.

This chapter reads the entry-level result and the support-agent compression result as the same possible mechanism seen from two ends: when AI makes codified expertise easier to obtain, work whose value rests heavily on that expertise faces pressure. The studies do not establish that junior roles are generally the most codified, or that senior premiums will erode in this order; those are the book’s inference.

Bound the evidence itself: one firm, one occupation, a support context with unusually clean productivity metrics and unusually codifiable expertise. Software engineering is not customer support. What transfers is the hypothesis worth testing, not the 34%.

A newer result complicates the compression story without removing it. In a pre-registered three-month trial with 133 patent lawyers, AI assistance raised drafting quality (+0.34 SD at 10 days, +0.38 SD at 90 days, larger for juniors) — but durable unaided judgment afterwards concentrated in seniors (+0.45 SD), while juniors bifurcated rather than improving on average (Autor et al., NBER w35720, 2026).

That is NBER working-paper evidence, one occupation, Google-funded: it supports, within those conditions, that foundational expertise may be a prerequisite for extracting lasting skill from AI-assisted practice. Compression of immediate output and differentiated learning can both be true — and the ladder problem below gets harder, not easier, if juniors get the output gain without the judgment gain.

Where the ground is solid: the jagged edge

If the codified part is being flattened, what is not?

Dell’Acqua and colleagues ran a preregistered field experiment with 758 BCG consultants, roughly 7% of the firm’s individual-contributor consultants. On 18 realistic tasks chosen to sit inside current AI capability, consultants using AI completed 12.2% more tasks, 25.1% faster, at significantly higher quality. On a complex task deliberately chosen to sit outside that capability, consultants using AI were 19% less likely to reach a correct solution than those without it — because they extended unwarranted trust to confidently-presented, substantively wrong output (Dell’Acqua et al., 2023).

The authors named the shape of the problem: a jagged technological frontier. Some tasks fall easily inside current capability. Others, apparently similar in difficulty, fall outside it. The boundary is irregular and it is not visible from the task description.

So here is the job that has value: knowing where the edge is. The important property is not that the frontier is impossible to codify. It is that it cannot be learned once and banked. Any rule about where models work is dated by the capability, task, context, and verifier that produced it, so frontier judgment is a standing obligation to remeasure. On the outside-frontier task in this study, failing to respect that boundary made AI-assisted participants less likely to reach the correct answer than the control group.

That is the first of four responsibilities this book keeps outside the model, and Chapter 3 turns it into an engineering rule rather than an intuition.

The four jobs this book keeps outside the model

Chapter 1 drew the applied arrangement: the model supplies a proposal, the process holds state and checks it, and the person supplies intent and retains authority. That was an architecture picture. Read it again as a map of where to stand.

The jobWhat it isWhy this book keeps it outside the model
IntentDeciding what should be done, and what would count as doneThe model can propose objectives, but the objective and success condition have to be owned outside the generator if its output is to be judged against them.
AuthorityDeciding who is permitted to cause which effectsCapability is not authority (Chapter 20). Permission must come from an explicit actor or policy rather than from the component proposing the action.
VerificationObtaining evidence about whether the target state holdsThe generator’s assertion is not evidence that its proposal worked. Verification must be independent of generation: mechanical where an adequate oracle exists, otherwise grounded in sources, observed state, external outcomes, or human judgment (Chapter 21).
Frontier judgmentDeciding when the stochastic component is the right mechanism at allThe frontier is jagged and moving. This book therefore keeps the decision to use AI as a standing human responsibility that is revisited as measured capability changes.

These four responsibilities are not claimed to be immune to automation. The architectural claim is narrower: improving generation does not remove the need for an objective, permission to cause effects, evidence about outcomes, or a decision about whether generation is the right mechanism. As generation becomes cheaper and more abundant, the process has more proposals to aim, authorize, and verify.

That is what it means to not stand in front of the steamroller. Not “move into management,” which is advice about a job title. It means occupying the part of the process whose demand is created by the thing doing the flattening.

Here the career argument and the architecture argument turn out to be one argument. Every remaining chapter of this book builds machinery around these responsibilities: explicit acceptance criteria and selected context so intent is legible (Chapters 14–15), a ledger and artifact store so what happened survives (Chapters 16–17), claims and evidence levels so support stays distinguishable from assertion (Chapter 18), capability kept separate from authority (Chapter 20), independent verification (Chapter 21), and a small deterministic scheduler that chooses the next operation before model selection becomes relevant (Chapter 28).

Applied AI is not a set of tricks for getting more out of a chat window. It is the engineering discipline of the position you want to be standing in.

The same responsibilities decide something closer to home. Once an assistant can build software, deciding what your own tools should do, what they may change, and how you would know a change helped is intent, authority and verification applied to your working environment. The career inference is deliberately weaker than “if AI can do it, everyone can”: when a task is codified, repeatable, and cheaply checkable, assume its automation boundary can move and keep measuring where your distinctive responsibility still lies. Chapter 30 returns to that.

Where this argument is weakest

A chapter this sure of its main claim owes you its own best counter-arguments. Here are the ones I find genuinely hard.

The ladder problem is real and this chapter does not solve it. The route to senior judgment has historically run through the junior codified work, where you learned where the frontier was by being wrong about it on small things for several years. If the steamroller removes the bottom rungs — and “reduced hiring rather than separations” says precisely that it does — then “acquire senior judgment” is not actionable advice for someone who cannot get hired to acquire it. I do not have a good answer, because this is a structural problem requiring a structural response, and individuals absorbing it as a personal failure are misreading it.

The composition problem. Advice that works for one person can be arithmetically impossible for a cohort, since if everyone moves to directing AI, the ratio of directors to directed work becomes absurd. “Get out of the way of the steamroller” describes a smaller destination than the road it is flattening, and anyone telling you otherwise is selling something.

Verification might automate further than I am betting. This book’s central wager is that verification resists automation because it requires contact with reality, which is a bet rather than a theorem. Automated test generation, formal methods, and simulation all push on it, and if verification collapses into the model, a good part of this chapter’s advice collapses with it.

The evidence is young and mostly descriptive. The strongest labor result here is explicitly labeled by its authors as descriptive rather than causal, attenuates under education controls, and runs larger in its sample than in national benchmarks, while the productivity results come from single firms or single occupations. The METR reversal described in the next section is a live demonstration that confident readings of this literature can age badly within months.

The metaphor invites fatalism. “Steamroller” can be read as “nothing you do matters.” That is the opposite of the point. The point is that where you stand is a decision, it is available to you, and the evidence about which ground is solid is better than it was two years ago.

Being wrong responsibly

This book will be wrong about things.

Not as a disclaimer. As a statement about what kind of document this is. Everything here is a view from one position on a hill that is still being climbed. From where we are standing in late 2026, certain shapes are visible and certain ones are not, and some of what currently looks like a permanent feature of the landscape will turn out to have been a cloud.

The useful question is not whether a book about AI will age badly. It is whether it ages badly in a way you can detect and correct, or in a way that quietly misleads you for two years.

So here is what being wrong responsibly actually looks like, from a lab that does this well.

In July 2025, METR published a randomized controlled trial on AI coding tools. Sixteen experienced open-source developers, 246 real tasks in their own repositories — mature projects averaging over a million lines of code — each task randomly assigned to allow or disallow AI assistance. The developers forecast beforehand that AI would cut their completion time by 24%. Afterwards, having done the work, they estimated it had cut their time by 20%. Measured, allowing AI increased completion time by 19% (METR, 2025).

That result traveled a long way, and deservedly: a 39-point gap between what skilled practitioners believed about their own productivity and what was measured is a serious finding.

Then, in February 2026, METR published an update saying their experiment design had a problem. Developers were increasingly declining to participate in the no-AI condition, even at $50/hour, and between 30% and 50% of participants were avoiding submitting the very tasks they expected AI to accelerate. That is selection bias pointing directly against the headline. Their newer estimates moved the other way — around −18% time for returning developers (confidence interval −38% to +9%) and −4% for newly recruited ones (−15% to +9%) — with both intervals crossing zero. METR’s own summary is that developers are likely more sped up in early 2026 than their early-2025 estimate suggested, and that the new data is only very weak evidence (METR, 2026).

Read what happened there. A careful lab ran a clean experiment, got a striking number, published it, kept looking, found the flaw themselves, and stated plainly how weak the replacement evidence is. Nothing was retracted and nothing was defended past its evidence.

That is the standard this book holds itself to, and it is why Part 5 spends five chapters on experiments, including one in which a result the author wanted to be true did not survive replication. The durable lesson from METR is not the 19% slowdown or the original 39-point gap between belief and measurement. The later design update changed what the experiment could support and moved the estimated effect toward the developers’ own expectations, with wide intervals crossing zero. The durable lesson is methodological: neither a striking estimate nor a confident self-assessment gets to grade itself; the design, selection process, and later evidence remain part of the result.

How to read the rest of this book

Claims in this book come in different strengths, and the book tries never to let a weaker one borrow the authority of a stronger one:

  • Measured — an empirical result produced by a study or experiment, reported with the population or task set, method, and relevant bounds.
  • Inspected — an observation about source code that was actually read, identified by file and symbol.
  • Argued — a position defended from stated premises, like the determinism rule in Chapter 3.
  • Predicted — a claim about how things will go. The steamroller’s future is one of these; its past, and its measured pace so far, are evidence for the prediction rather than the prediction itself. It is the weakest category, and everything in it should be held loosely.

In the early chapters the strength is stated in the prose. From Chapter 19 onward, where the book leans on its own experiments, claims about CodeAI also carry explicit tags that separate a pinned measurement from a demonstration that ran but was not pinned, a frozen historical report, code that was only read, and evidence that is still owed.

Where a number appears, it has a source and a stated limit. Where this book proposes a contract stronger than the code currently implements, it says so. Where an experiment failed, the failure is the chapter.

Do this now

Fifteen minutes, one page. Audit your own exposure.

  1. List the five things you were paid to do last month. Be concrete — tasks, not a job title.
  2. For each, mark whether its value is the codified part (knowing what is written down, executing a described procedure) or the tacit part (judgment about what should be done, or about whether it worked).
  3. For each codified one, write the sentence: if a model did this at 90% quality for a hundredth of the cost, what would still need me? If you cannot finish the sentence, that is the finding.
  4. Now mark which of this chapter’s four jobs — intent, authority, verification, frontier judgment — you actually hold, as opposed to being adjacent to.

To turn the audit into a number, score it as a constructed teaching sketch (illustrative tasks, not measured data). Mark each task C (value is the codified part) or T (value is tacit judgment), then count codified tasks where you hold none of the four jobs:

Example taskC / TFour-job held?Exposed?
Reset passwords per runbookCnoneyes
Triage routine support ticketsCnoneyes
Write weekly status summaryCintentno
Judge an ambiguous outage escalationTfrontier judgmentno
Sign off a production migrationTauthority, verificationno

Here 2 of 5 tasks are exposed (2 codified tasks with no job held). You can now do one new thing: compute your own exposed count and name the boundary — any task scoring C with no job held is the section of road to move off first, before any career decision is made on vibes.

Keep the page. Chapter 7 asks you to measure systems; this exercise is different. It is a constructed self-audit in which you are both the instrument and the subject, not an empirical measurement of career risk.

Failure modes

  • Reading “no economy-wide displacement” as “nothing is happening.” Both facts come from the same paper. The aggregate is calm; the 22–25 cohort in exposed occupations is not.
  • Reading the entry-level data as a reason to disparage juniors. The mechanism is reduced hiring in codified roles, not a judgment about the people who would have filled them.
  • Trusting confident output near the frontier. On the outside-frontier task in the BCG experiment, AI-assisted participants were less likely to reach the correct answer than the control group.
  • Treating your own productivity estimate as measurement. The original METR trial reported a 39-point gap between developers’ expectations and its measured estimate; the later design update changed the estimate substantially. Self-assessment is evidence about belief, not a substitute for outcome measurement.
  • Treating “move up the stack” as a completed action. Frontier judgment is a standing obligation to re-check, because the frontier moves.
  • Quoting any number in this chapter without its bound. Every one of them is from a specific sample in a specific period.

What this chapter established

  • The historical premise is established: automation has displaced codified parts of skilled jobs before. The current AI claim is dated and narrower: measured model capability on software tasks has been moving quickly, so codified, checkable work is the part whose automation boundary deserves the closest watch.
  • Once a codified part can be automated cheaply and adequately, excellence at that part alone is not durable protection. The practical question is which responsibilities remain when that component becomes cheap.
  • The steamroller is a useful model — tacit → codified → automated → commodity — not a law that every codified activity inevitably completes the sequence.
  • Measured: no economy-wide displacement, but a 19% kept-pace shortfall for 22–25-year-olds in AI-exposed occupations, driven primarily by reduced hiring and concentrated where AI substitutes rather than complements; the authors label these patterns descriptive rather than causal.
  • In one customer-support setting, AI assistance produced much larger immediate productivity gains for novice and lower-skilled workers than for experienced workers. This chapter treats compression of codified expertise as a hypothesis that result motivates, not a universal labor-market finding.
  • The frontier is jagged. On the deliberately outside-frontier task in the BCG experiment, AI-assisted participants were less likely to reach the correct answer than the control group.
  • Four responsibilities this book keeps outside the model: intent, authority, verification, and frontier judgment. The claim is architectural: stronger generation does not remove the need for them.
  • The honest weaknesses: the broken ladder, the composition problem, the possibility that more of verification automates, and evidence that is young and often descriptive.
  • This is a dated position, and being wrong responsibly means retaining the correction with the finding — the standard the chapter draws from METR’s 2025 result and 2026 design update.

Next

Three of those four jobs — intent, authority, verification — are engineering problems, and the rest of this book builds them. The fourth, frontier judgment, needs a rule before it can be built into anything, because “use AI where it’s good” is not something you can put in a file.

The next chapter turns it into one. It argues that a component belongs on the deterministic side unless it demonstrably cannot be, that doubt about which side is itself the answer, and that the stochastic box is more stochastic than its configuration claims — temperature=0 does not make a model deterministic, and the reason has nothing to do with sampling.

Continue with If There’s Any Doubt, It’s Deterministic.

References