A Revolver, Not a Foundation
The model underneath your system is a replaceable occupant, and even a stable product name can hide behavioral drift. Finish the frame; keep turning the chamber. Treat releases as experiments, and turn preserved, still-valid work into a regression corpus before you swap.
Part 2 — Get the Model Out of the Chat Box
The model will change; the frame should survive it
The book you could write better next year
Write a book with AI this year and you will be able to write the same book better next year. And better again the year after. The same is true of the code, the research synthesis, the review process, the classifier — anything where a model does part of the work.
That is good news. It is also a statement about the ground you are building on, because the component supplying stochastic capability does not behave like an ordinary stable contract.
Databases and compilers can change and regress too. The difference is that ordinary software is normally built against explicit semantics that upgrades are expected to preserve. With a hosted model, the API can remain compatible while the distribution of useful answers, failures and instruction-following changes materially.
The model underneath an AI system should therefore be treated as a replaceable occupant rather than as the foundation itself. The evidence below establishes rapid movement on one capability measure, behavioral drift under a stable product name, and negative flips across model updates. It does not establish the cadence or magnitude of future releases. Design for replacement because replacement and drift occur, not because a trend line guarantees what comes next.
What do you build on, when an important component underneath you may be replaced by something behaviorally different while the process still has to survive?
The ground is moving, measurably
Three pieces of evidence, each describing a different way it moves.
Frontier capability has moved fast on this measure. METR proposed a practical measure: the length of task, in human time, that a model can complete with 50% success. On their software and research task suites, the frontier’s 50% time horizon doubled roughly every seven months from 2019, possibly faster in 2024, with Claude 3.7 Sonnet at around fifty minutes when they published (Kwa et al., 2025). A seven-month doubling compounds to roughly ten times every two years. That is the arithmetic of the fitted historical trend, not a forecast or a general rate for every kind of work. The authors explicitly limit its external validity and make extrapolation conditional.
The “same” model moves too, in no particular direction. Chen, Zaharia, and Zou compared the March and June 2023 versions of GPT-3.5 and GPT-4 — same product names — across seven task families. GPT-4’s accuracy at distinguishing prime from composite numbers fell from 84% to 51%, associated with a drop in how well it followed chain-of-thought prompting; code-generation formatting errors increased for both models; some other capabilities improved (Chen, Zaharia & Zou, 2023). Their conclusion is the one to keep: the behavior of the “same” LLM service can change substantially in a short time, so it needs continuous monitoring.
A small note that is itself on-theme. An earlier version of that paper reported a far more dramatic prime-number collapse, and that larger figure is still the one circulating. The authors revised it. Cite the current version.
Upgrades break specific things while improving the average. Echterhoff and colleagues studied what happens to task-specific adapters when the underlying base model is updated. They found negative flips — instances the old model got right that the new one gets wrong — across a diverse range of tasks and models, even when the downstream training procedure stayed identical. Their compatibility method cut negative flips by up to 40% when moving from Llama 1 to Llama 2 (Echterhoff et al., 2024). The method is not the point. The point is that “the new model scores higher” and “the new model still does the thing I depended on” are different claims, and the first does not imply the second.
So: capability compounds, named models drift, and upgrades regress on individual cases. None of those is a foundation.
Finish the frame, turn the chamber
Chapter 8 argued that bounded work should be able to finish. This chapter makes a different claim: the model occupying a stochastic role should remain replaceable after that work finishes.
Chapter 3’s rule was not that every component outside generation freezes forever. It was to keep operations deterministic where explicit state and rules can compute them, and to confine stochastic proposals where they are actually needed. The frame in this chapter means that owned process machinery: identities, records, context rules, deterministic checks where adequate specifications exist, authority and policy seams, and adapter contracts.
You finish the frame. You keep turning the chamber.
“Finish” does not mean immutable. Specifications change, checks improve, policies evolve and provider integrations sometimes need work. It means those changes are driven by the process’s requirements rather than by model behavior leaking through every layer.
A chamber is the replaceable model role inside that frame. If the boundary holds, changing its occupant becomes a bounded re-evaluation plus whatever configuration or adapter work the new occupant requires, rather than an architectural rewrite.
The frame is also the part you own. Your context rules, records, checks and policies can survive occupants coming and going. That is what lets a working environment built around you outlast a particular model or vendor — the argument Chapter 30 closes the book with.
Spread model calls through parsing, policy, routing and checking, and the model-dependent regression surface grows. A release can then force re-evaluation across many unrelated mechanisms. That is a maintenance cost this architecture is designed to avoid; it is not evidence that such teams literally cannot finish projects.
The Chapter 3 benefit is narrower and stronger: every operation made genuinely deterministic no longer depends directly on the behavior of the current model occupant, even though its own specification and dependencies can still change.
The frame holds still while the chambers turn:
flowchart TD
subgraph FRAME["deterministic frame — finished"]
direction TB
SP["specification"]
CX["context assembly"]
VG["verifiers"]
LP["ledger · policy · router"]
end
subgraph CH["chambers — one stochastic slot each"]
direction TB
C1[("fast-classify<br/><i>occupant, version, price, date</i>")]
C2[("deep-review<br/><i>occupant, version, price, date</i>")]
end
FRAME --> CH
CH -.->|"new release: re-run<br/>against history, not rewrite"| CH
The chambers
A revolver holds several rounds and you choose which one to fire. The useful version of that picture: your system has several chambers — logical roles — and each is occupied, at a given moment, by a specific model.
A chamber is named for its job, not its occupant: fast-classify, draft, deep-review, critic-a, critic-b. The design requires the concrete occupant to be recoverable when work is executed. Do not make configuration claim provenance it does not have: CodeAI’s current chamber mapping stores the logical name, adapter, model and route information; it does not itself store an observed provider revision, price or load date. Chapters 11 and 13 attach execution and pricing provenance separately, with unknown values left unknown.
What you choose on, per chamber:
| Property | The question | How you measure it |
|---|---|---|
| Speed | Is it fast enough for where it sits? | Latency percentiles per attempt and call, including retries |
| Quality | Does it satisfy this chamber’s declared criterion? | Frozen-task outcomes with uncertainty; accepted and correct kept separate where independent correctness exists |
| Knowledge | Does it know what this job requires — domain, recency? | Targeted items in the frozen set; do not infer a training cutoff that was not actually observed |
| Applicability | Does it fit the job’s shape — context length, modality, tools, structured output? | Mechanical capability and interface constraints |
| Utility | Is the measured outcome worth the process cost here? | Cost per accepted outcome, with wrong acceptances reported; cost per correct acceptance where correctness is independently known |
| Gain | Is it better than the current occupant on this chamber’s work? | Paired per-item difference against the incumbent, plus flips in both directions under a frozen decision rule |
The last row decides swaps, and it is chamber-specific. A new model can win deep-review and lose fast-classify in the same week. A global “upgrade to the new model” decision throws that information away.
An illustrative swap shows what the decision looks like on paper — not a measured result, a worked example of the arithmetic. Suppose one chamber’s frozen set has 40 items. The incumbent passes 28, the candidate passes 31. Paired per item: 26 pass on both, 7 fail on both, 5 pass only on the candidate, 2 pass only on the incumbent.
The paired difference is +3 items (+7.5 points) with a standard error near 6.5 points, so the interval comfortably crosses zero — a higher average that could easily be chance, with 2 negative flips on items you may depend on. Price it: say the replay cost $0.40 on the incumbent and $1.20 on the candidate, so cost per passing item is about $0.014 against $0.039. The candidate wins the average and loses utility by nearly three to one.
That is not a swap; it is at most a note saying which five items improved and whether the two flips matter, written down with both occupant versions, the numbers, and who approved.
This is not an invented abstraction. CodeAI resolves a chamber name to its occupant through .codeai/config.toml, and the mapping is small enough to read at a glance. From one of CodeAI’s regression tests:
[models.deep-review] # the chamber, named for its job
adapter = "opencode"
model = "mimo-v2.5" # the current occupant
protocol = "chat_completions" # how that occupant is served
ModelConfig.resolve_chamber("deep-review") returns that mapping, and nothing else in the system needs to know which model it names. Changing the occupant is an edit to configuration, not to code.
The provider layer adds a versioned pricing table that returns no cost at all — never zero — for a model it does not recognize (PRICING_VERSION and estimate_cost_usd in providers.py). Chapter 28 finds the limit of that caution: the same table matches model names by prefix, so gpt-4o-mini was priced at gpt-4o rates. Refusing to invent a price for an unknown model is not the same as pricing a known one correctly.
That is the chamber mechanism in its smallest form: a name the rest of the system depends on, bound to an occupant that can change. The full release protocol below is never built in this book as a single mechanism. Its parts appear separately: Part 5 builds paired comparisons under frozen decision rules, and Chapter 28 puts several occupants on the named rungs of an escalation ladder.
The price-performance frontier moves too
The evidence is stronger for one economic trend than for the claim that every new frontier launch costs more than what it replaces. Epoch AI measured the price of reaching fixed benchmark-performance milestones falling at rates from 9× to 900× per year depending on the milestone, while cautioning that the fastest recent declines may not persist (Epoch AI, 2025; see Chapter 6).
That gives the revolver a more durable economic logic:
- Re-evaluate a chamber when a new model, a meaningful price change, or a change in availability or terms could alter the decision.
- Prefer the least costly permitted occupant that satisfies the chamber’s frozen promotion rule; a lower bill does not compensate for an accepted-but-wrong outcome the rule forbids.
- Keep a more expensive occupant only where the measured additional value justifies the additional process cost.
Keep that up and you accumulate something no vendor can sell you: a chamber-specific record of quality, failure shape, latency and cost on your own work. That map is where the advantage lives.
A release is an experiment, not an upgrade
When a new model arrives, the tempting move is to try it on a few memorable cases, like what you see, and switch. Chapter 7 already gave the correction: sample size and repetition come from the effect you need to detect, the variance in the process and the decision you intend to make — not from habit or convenience.
The protocol instead, using what Part 1 already built:
- Pick one chamber. Never evaluate “the new model” in the abstract.
- Define the treatment before the run. If you want to isolate the model change, hold the prompt and other controllable inputs fixed. If each occupant needs its own tuned prompt, compare the deployable model-plus-prompt bundle and label the treatment that way; do not attribute the whole difference to the model.
- Replay the frozen task set with enough repetition for the question you are asking. A few draws can expose obvious run-to-run variation, but
n ≥ 3is not a general statistical rule. Freeze the repetition and analysis plan before seeing the result. - Compare paired, per item, against the incumbent. Report the paired difference with appropriate uncertainty and count flips in both directions. Critical regressions belong in the promotion rule before the run; a higher average does not erase them automatically. And a win on an inadequate verifier is not a win.
- Keep the reference evaluation fixed. If a critic is part of it, do not swap that critic in the same experiment as the occupant being judged. Otherwise the treatment changed twice.
- Price the whole result. Report cost per accepted outcome alongside accepted-but-wrong outcomes; where independent correctness is available, also compute cost per correct acceptance.
- Decide per chamber and record the decision — date, treatment identities, versions where observed, measurements, promotion rule and approver — and keep a tested rollback path where the provider still permits one.
That is a deliberate release evaluation rather than an upgrade by reputation.
If an occupant is reached through a mutable provider alias or another route whose behavior can change without a configuration edit, scheduled re-evaluation belongs in the same machinery. Exact immutable snapshots, where they genuinely exist, may justify a different monitoring cadence; the policy should record which case you are in.
Your history can become a regression suite
Here the durable record from Chapter 9 pays off again once the later construction chapters add exact call identity, preserved observations and context provenance.
Past work is not automatically a regression test merely because it was recorded. It becomes a useful regression case when you can recover the task, the relevant input and context identity, the criterion or verifier in force, and enough outcome evidence to know what property the previous run established. An earlier accepted output is evidence about the old process; it is not ground truth merely because it survived it.
With those pieces intact, preserved work can seed a regression corpus for a candidate occupant: does the new treatment still satisfy the same still-valid criterion on this case, and where do flips occur?
It can also seed a regeneration queue where regeneration is actually wanted. A chapter, review or other artifact can be proposed again against its recorded objective and then compared under the current evaluation before anything replaces the existing artifact. That is not “the new model is better, therefore rewrite everything.” The objective, context assumptions and evaluation all have to remain applicable.
Without preserved inputs and provenance, an exact historical replay may be impossible. With them, replay becomes an available experiment rather than a reconstruction from memory.
Two limits, both from Part 1. Regenerating everything is Chapter 6’s invoice in a new costume; regenerate where measured gain justifies the process cost. And not everything should revolve. Work that has been distilled into deterministic machinery, or a stable artifact whose relevant criteria remain satisfied, does not need a new model simply because one exists.
Where this breaks
- Trends end. The seven-month doubling is a historical fit with a stated external-validity caveat. Build for replacement because releases happen, not because a curve promises the next one will be better.
- Prompts are partly model-specific. A prompt tuned for one occupant can underperform on another, so a swap may need prompt work. The specification should record the prompt per chamber per occupant.
- Vendors set part of the schedule. Deprecations force swaps whether or not a candidate has been evaluated. The protocol above is how you avoid being forced into an untested one.
- Evaluation costs money every release. The harness is not free to run (Chapter 7), which is another reason to evaluate the chambers that matter rather than everything.
- Drift without a swap is still drift. A pinned model name does not guarantee pinned behavior.
Do this now
Twenty minutes. Load the cylinder.
- List every place your system calls a model. Give each a chamber name for its job, not for the model in it.
- For each chamber, record the current occupant: provider, exact model id, version or snapshot date, input and output price, and the date you last evaluated it.
- Count how many places in your code a concrete model id is written directly rather than resolved from a logical name. Each one is a place where a release costs you a code change instead of a configuration change.
- For your most important chamber, write down what it would take to replay its frozen task set against a candidate tomorrow: which tasks, which checks, which critic, and how you would compare per item.
If step 4 has no answer, that — not a model choice — is your next piece of work.
Failure modes
- Treating the current occupant as a foundation. The architecture now depends on one model identity and failure shape remaining stable.
- Hard-coding model ids. Every occupant change becomes a code change instead of a configuration and evaluation decision.
- Upgrading globally. A candidate can improve one chamber and regress another; one global switch discards that distinction.
- Adopting on aggregate score. Averages can hide negative flips on cases the process depends on.
- Changing the evaluator and the candidate together. The treatment changed twice, so the observed difference cannot be attributed to the occupant alone.
- Assuming a stable product name means stable behavior. Provider aliases and hosted services can drift; record whether the identity you hold is actually immutable.
- Spreading model-dependent work through the frame. Each occupant change then expands the regression surface into parsing, policy, routing or checking that did not need to depend on the model.
- Regenerating by reflex. Re-run where measured gain justifies process cost; leave stable work alone when its relevant criteria still hold.
- Treating history as gold. A preserved accepted output is not automatically a valid regression oracle; preserve the task, criterion, provenance and outcome evidence too.
What this chapter established
- The current model occupant is not a foundation. METR measured a roughly seven-month historical doubling in frontier task horizon on its task distribution, Chen and colleagues measured substantial behavior change under the same GPT product names, and Echterhoff and colleagues measured negative flips across base-model updates. Each result is bounded to its study; together they justify designing for replacement and regression measurement.
- Chapter 8 and this chapter describe different parts of a system: finish the frame, turn the chamber. The frame is owned process machinery that can evolve without being coupled to every model release; the chamber is the replaceable model role inside it.
- Keeping model-dependent behavior behind explicit seams reduces the surface that must be re-evaluated when an occupant changes. It does not make the rest of the software immutable.
- Chambers are named by job and evaluated separately. CodeAI’s current logical-model mapping is the smallest implemented form: a chamber resolves to adapter, model and route configuration, while observed revision, attempt and pricing provenance belong to recorded execution.
- The measured economic trend is that fixed benchmark capability has become dramatically cheaper, at very different rates across milestones. Re-evaluate chambers when capability, price, availability or terms move; do not assume every new frontier launch has the same price pattern.
- A release is an experiment: define one chamber and one treatment, freeze the evaluation design, compare paired item-level outcomes with appropriate uncertainty, preserve flips, keep the reference evaluation fixed, measure process cost and apply a promotion rule written before the result.
- Preserved history can seed a regression corpus and regeneration queue when the task, inputs, criterion, state assumptions and outcome evidence remain valid. An old acceptance is not ground truth by itself.
- The durable advantage is a measured chamber-specific record of which occupants earn their place on your work, under which criteria, at what cost and with which regressions.
Next
Every mechanism in this chapter rests on one unglamorous fact being true at the lowest level: each call records exactly which model answered it. Which provider, which model id, which snapshot, at what price, after how many attempts, with what raw output.
If the call does not record that, chambers cannot be compared, negative flips cannot be counted, drift cannot be detected, and your history is not a regression suite — it is a pile of text with no provenance.
So construction starts there.
Continue with The Smallest Useful Model Call.
References
- Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, et al. (METR). Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499, 2025. https://arxiv.org/abs/2503.14499
- Lingjiao Chen, Matei Zaharia, and James Zou. How Is ChatGPT’s Behavior Changing over Time? arXiv:2307.09009 (v3), 2023. https://arxiv.org/abs/2307.09009
- Jessica Echterhoff, Fartash Faghri, Raviteja Vemulapalli, Ting-Yao Hu, Chun-Liang Li, Oncel Tuzel, and Hadi Pouransari. MUSCLE: A Model Update Strategy for Compatible LLM Evolution. Findings of the Association for Computational Linguistics: EMNLP 2024. https://arxiv.org/abs/2407.09435
- Epoch AI. LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks. Data Insight, March 12, 2025. https://epoch.ai/data-insights/llm-inference-price-trends
Implementation sources: CodeAI, src/codeai/modelconfig.py (ModelMapping, ModelConfig.resolve, ModelConfig.resolve_chamber, load_model_config) and src/codeai/providers.py (PRICING_VERSION, PRICING_TABLE, estimate_cost_usd). Inspected, not yet exercised in a chamber-swap walkthrough.