What Should Happen Next?
Which operation is needed before choosing a model? The deterministic scheduler decides the operation, a measured execution ladder decides escalation, and a model-router challenger waits unrun. The ladder was cheaper per accepted outcome and produced one more correct acceptance — and still failed its frozen adoption rule on one accepted-but-wrong answer.
Part 6 — Put Intelligence Into the Process
The pieces are all built. This part puts them under one policy and then joins them.
It starts with the decision that usually gets skipped: which operation does the process need next — a model call, a deterministic check, a person, or a stop? That choice comes before any question about which model, and this part measures what climbing from cheap to expensive actually bought. Then one task runs through every boundary the book has built, from intent and authority to effect, verification and acceptance, to find out whether mechanisms that each work alone still hold where they hand off to one another.
The last chapter changes the subject from the process to the person using it. When software can be built around one individual, this architecture is what lets a tool shaped around your own work keep improving without losing its discipline.
Choose the operation before you choose the model
When an AI process is not getting anywhere, the reflex is to reach for a better model. That reflex answers a question nobody asked. The real one is what kind of operation the situation needs, and its honest answers include running a deterministic check, asking a person, and stopping.
Three decisions collapse into the phrase “route to the right model”:
which operation → which rung → which model
They fail differently. A cheaper model cannot repair choosing the wrong operation. A stronger model cannot repair a missing check. And a decision procedure that itself calls a model has added a stochastic dependency to the part of the system that was supposed to govern the others.
By the end of this chapter you will be able to separate those decisions in your own process: a deterministic policy that names the next operation without asking a model, an escalation ladder that records why each rung declined, and an honest account of which decision your evidence actually covers. The measured part here is escalation, and it is also the chapter’s cautionary tale: the cheaper ladder won on cost, produced one more correct outcome, and still failed its own adoption rule, because it accepted one wrong answer.
A claim awaiting a check
Picture a patch sitting open with one review comment unresolved: a retry test that may or may not cover the new branch. Three things could happen next. A model could propose a fix. A deterministic check could run the test suite. A person could be asked whether the branch matters at all. Each costs something different, fails differently, and answers a different question. “Try a stronger model” is not on that list until something has established that generation is the missing operation.
Chapter 27 ended with a policy that had to earn promotion and did not. Every experiment in Part 5, though, spent model calls freely in order to ask its question. This chapter turns the same discipline on the spending itself: the next model call should have a reason, and the model stops being the implicit controller of the loop.
Which operation is needed before choosing a model?
Three decisions that usually get merged
“Route to the right model” hides three separate decisions:
| Decision | The question | CodeAI mechanism | Evidence in this chapter |
|---|---|---|---|
| Operation | Call, check, ask a person, or stop? | decide_next_step, a pure function | Source inspection and current seam tests; a historical matrix demo of an older version |
| Escalation | This rung declined — climb, or go to a person? | run_ladder, with recorded decline reasons | Stage 29B, measured |
| Selector | Should the operation decision itself contain a model call? | Router contract and model-based router | Designed with a falsifier; unrun |
Keeping them apart changes what failure means. A failed model call is not automatically a reason to buy a stronger one; it may be a reason to run a check, or to stop climbing and ask a person. And the evidence for one decision does not transfer to another. Stage 29B — this chapter’s live experiment, in which forty extraction items climb two different escalation arms against real providers — tests escalation. It says nothing directly about whether an operation selector should be deterministic — which is the question the unrun experiment exists to ask.
The literature has names for two of the halves. Modular architectures that route among language models, knowledge sources, and discrete reasoners treat which component answers as a decision separate from generation (Karpas et al., 2022). That supports treating operation choice as its own layer, though a modular diagram does not demonstrate a fixed-priority scheduler, and the mapping onto CodeAI is the book’s. Learned routers attack the neighboring decision: given that a model will be called, which one, trading quality against cost with routers trained on preference data (Ong et al., 2024). Neither paper validates CodeAI’s ladder.
The operation seam
Here is the whole current policy, from CodeAI source:
def decide_next_step(query: SchedulerInput) -> SchedulerDecision:
if query.process_complete:
return SchedulerDecision(Operation.STOP, "the process is already complete", POLICY_VERSION)
if query.process_budget_exhausted:
return SchedulerDecision(Operation.STOP, "process budget exhausted", POLICY_VERSION)
if query.unresolved_effect:
return SchedulerDecision(
Operation.ASK_HUMAN, "unresolved_effect_requires_reconciliation", POLICY_VERSION
)
if query.has_required_verification:
return SchedulerDecision(Operation.CHECK, "required deterministic verification exists", POLICY_VERSION)
if query.requests_independent_proposals and not query.model_budget_exhausted:
return SchedulerDecision(Operation.CALL, "task requests independent proposals", POLICY_VERSION)
if query.requires_human_authority_for_next_effect:
reason = "next effect requires human authority"
if query.model_budget_exhausted and query.requests_independent_proposals:
reason += "; model budget blocks proposals"
return SchedulerDecision(Operation.ASK_HUMAN, reason, POLICY_VERSION)
if query.model_budget_exhausted and query.requests_independent_proposals:
return SchedulerDecision(Operation.STOP, "model budget blocks proposals", POLICY_VERSION)
return SchedulerDecision(Operation.STOP, "no epistemic operation required", POLICY_VERSION)
No model, no I/O, no ledger write. Policy epistemic-v3 orders seven facts, and its most important design choice is the one that is easiest to miss: two different budgets. Model-budget exhaustion blocks CALL and nothing else. Process-budget exhaustion stops everything. Collapse them into one “budget exhausted” flag and a system that has spent its cognition allowance can no longer run the check it still owes.
Two of the seven arrived later and are worth naming. A process the record shows as complete stops, rather than being re-offered work. And an unresolved effect — an action whose outcome the record cannot settle, as Chapter 19 established earlier — outranks every ordinary next step and asks for a person, because reconciling it is a judgment about the world rather than about the ledger.
The policy as a decision tree, in precedence order — operation first, model choice later and elsewhere:
flowchart TD
IN["scheduler input<br/><i>five core flags shown here</i>"] --> P{"process budget<br/>exhausted?"}
P -->|"yes"| S1["STOP<br/><i>dominates everything</i>"]
P -->|"no"| C{"required verification<br/>exists?"}
C -->|"yes"| CK["CHECK<br/><i>survives a spent model budget</i>"]
C -->|"no"| R{"proposals requested<br/>+ model budget left?"}
R -->|"yes"| CL["CALL"]
R -->|"no"| H{"next effect needs<br/>human authority?"}
H -->|"yes"| AH["ASK_HUMAN"]
H -->|"no"| M{"model budget<br/>blocks proposals?"}
M -->|"yes"| S2["STOP<br/><i>model budget blocks proposals</i>"]
M -->|"no"| S3["STOP<br/><i>no epistemic operation required</i>"]
AC["ACTION<br/><i>in the enum; no input selects it</i>"]
style AC stroke-dasharray: 4 4
Ahead of the process-budget test sit the two later facts: a complete process stops, and an unresolved effect asks for a person.
Constructed calls against that function show the precedence, executed as teaching code rather than as a stage: 2
decide_next_step(SchedulerInput(
has_required_verification=True,
requests_independent_proposals=True,
requires_human_authority_for_next_effect=True,
process_budget_exhausted=True,
model_budget_exhausted=True,
)).operation # Operation.STOP: process stop dominates the other four flags
decide_next_step(SchedulerInput(
has_required_verification=True, model_budget_exhausted=True,
)).operation # Operation.CHECK: the check survives the spent model budget
Three boundaries are part of the design, not omissions to smooth over.
- ACTION is unreachable. The enum contains it; no input selects it. Effects are authorized downstream under explicit grants (Chapters 20 and 22). A seam path to ACTION would complete the enum cosmetically while bypassing that authority story. CALL and CHECK schedule work; neither authorizes an effect. 3
- The function decides; it does not know. It still consumes facts rather than gathering them, and that separation is deliberate: it is what lets the same state be replayed against a different policy version. What changed is where the facts come from.
ProcessStatenow projects them from the ledger — proposals, owed checks, budgets aslimit/consumed/known, unresolved effects, the directive’s effective grant — and refuses to invent what the record cannot support: an unknown budget is never “exhausted”, declared criteria are not an owed check until something exists to check, and atask.completedwith no acceptance behind it completes nothing. A caller may still hand the raw function any flags it likes; that call records nothing, and nothing downstream will accept it as a decision. Two current semantics are narrower than their names suggest:requests_independent_proposalsis derived from there being no successful proposal yet, whileindependent_proposal_countis projected but not consumed by the policy; and the projected human-gate flag is derived from pending acceptance without ACCEPT authority, not from a general forecast of every possible future effect. 4 - Decisions are recorded, with what they rested on.
decide_next_for_taskprojects, decides, and appends the decision with its reason, policy version, the full state snapshot, that snapshot’s digest, and the events behind it. A decision can therefore be reconstructed later rather than re-derived from whatever is true then — and it stays what it was: a decision recorded as CHECK still reads CHECK after the world moves, and replaying its recorded state through today’s policy still yields CHECK. 4
Deciding is not permitting, and neither is doing
A recorded decision is still only a decision. For a while that was literally all it was: the scheduler decided CHECK, a caller executed a write, and the two merely coexisted in the ledger with no reference in either direction. An operation with no decision at all ran identically. 5
Three questions, three answers, and they must not be merged:
decision what kind of work should happen next the scheduler
authority whether this actor may perform it the recorded directive chain (Ch 20)
execution what was attempted, and what was observed the operation itself (Ch 19)
The runtime now binds the first to what follows it, on a deliberately narrow rule: a process decision must never be silently bypassed while the resulting operation still appears to belong to that governed process. Two shapes stay legitimate, and the record tells them apart:
| Shape | What it must carry | What the runtime checks |
|---|---|---|
| Scheduler-governed | The decision that selected it | It is recorded, belongs to this task, selected this operation class, and still refers to the state it was derived from |
| External or manual | Its source | Nothing — but it is recorded as ungoverned, and it cannot pass for governed |
Claiming the scheduler’s name without naming a decision is refused; that masquerade is the whole point of the rule. Governance runs before authority and answers a different question, so an external action can be permitted by governance and still denied by the directive chain. 5
The freshness test is worth stating exactly, because it is coarse on purpose. A decision records the digest of the process state it was taken on; if that digest has moved, the decision no longer describes the world it was about and permits nothing. Any recorded operation for the task moves it — including an authority transition (Chapter 20), which expires outstanding decisions without either seam knowing about the other. This establishes that an operation was launched under a decision that still described the world at launch. It is not a lock, and between the check and the effect the world may move again. 5
And the unreachable ACTION now carries weight it did not before. Because no scheduler decision can ever select an effect, no effect can be scheduler-governed — an action naming a decision is refused outright. “A CHECK decision cannot silently become a WRITE action” holds not because the runtime compares them, but because no decision can license a write at all. Effects are gated by authority, exactly where Chapter 20 put them. 5
The historical scheduler matrix needs its version stamp attached. It ran all sixteen combinations of the older four-flag vocabulary against policy epistemic-v1: 16/16 parity between the code and a data-table interpreter, CALL/CHECK/ASK_HUMAN/STOP reachable, ACTION dead, and the same simultaneity precedence.
The committed policy has since become epistemic-v3: split budgets, plus the completion and unresolved-effect facts. Legacy spellings still migrate, and the frozen router corpus still validates, because its five-field states are explicitly accepted as a schema that predates the newer facts rather than being rewritten to fit them. No matrix covers the later rows. The later change to scheduler.py since that run is a dict-literal reformat with no behavioral change. 6 7
That demo also settled the stronger form of Chapter 1’s second bet, which promised to report the answer either way: that the router should be a specification rather than code. A five-row declarative table reproduced the if-chain exactly, but it duplicated the chain without fixing the real gaps — unpersisted decisions and precomputed flags. At this size, the book keeps the small function and drops the policy DSL. The weaker form of the bet stands: the decision path contains no model call. 6
The measured center: Stage 29B
The scheduler answers which operation. Stage 29B asks the question one level down: once the operation is “get an answer,” how should the process climb from cheap to expensive, and when should it stop climbing?
The workload. Forty synthetic cache-TTL extraction items with frozen gold: ten structured (CDN_TTL_SECONDS=86400), twelve prose, twelve with distractor durations, six with no single correct TTL.
The arms. The ladder climbs rule → free model → cheap model → strong model → person. Top-first sends every item to the strong model, then to a person.
The acceptance check. Every candidate passes through the same deterministic checker, blind to gold: valid JSON, a verbatim quote from the input, exactly one duration expression in that quote, and seconds equal to that duration.
The adoption rule, frozen before the run. The ladder is adopted only if all three hold at a $5-per-person-asked scenario: cheaper per accepted outcome; correct acceptances at least top-first’s minus one; accepted-but-wrong outcomes no more than top-first’s.
Calls went to named providers at declared prices, each arm-item pair in its own process, with every response body, receipt, and event preserved. 8
How the ladder climbs
run_ladder records every transition. A rung is entered with a reason, declines with one of five reason kinds, and the run ends resolved, with a person, or exhausted:
ladder.rung_entered rung, reason
ladder.rung_declined rung, reason ∈ { rule_declined:…, call_not_succeeded:…,
output_unavailable:…, check_failed:…,
check_error:… }
ladder.resolved | ladder.asked_human | ladder.exhausted
Two real traces from the preserved ledger show what that buys. A05 climbed one rung:
ladder.rung_entered rule "first rung"
ladder.rung_declined rule "rule_declined:0 ttl lines"
ladder.rung_entered zen-free "escalated from rule: rule_declined:0 ttl lines"
ladder.resolved zen-free
T07 climbed all the way, and its decline reasons never changed:
rule_declined:0 ttl lines
check_failed:quote_has_0_durations (free)
check_failed:quote_has_0_durations (cheap)
check_failed:quote_has_0_durations (strong)
→ ladder.asked_human
In the preserved Stage 29B revision, the report counted run_ladder at 61 lines and _check_decline at eight. That size claim is historical. More important, rung selection itself is deterministic: run_ladder advances through declared rungs from recorded decline reasons, while entering a model rung invokes the recorded model call in _candidate. The model supplies candidate cognition; it does not choose the next rung. 8 9
The result
Recomputed from the frozen analysis rows, matching the bundle’s tables: 8
| At $5 per person asked | Ladder | Top-first |
|---|---|---|
| Accepted outcomes | 33 (rule 10, free 23) | 31 (strong 31) |
| Correct acceptances | 32 | 31 |
| Accepted but wrong | 1 (A05) | 0 |
| People asked (of which: item had no single answer) | 7 (5) | 9 (6) |
| Model calls | 44 | 40 |
| Known model spend | $0.0056 | $0.0317, plus one call of unknown spend |
| Cost per accepted outcome | $1.061 | $1.453 |
| Summed model latency | 417 s | 315 s |
The preregistered output is recorded as ladder_justified: false, with the verdict label “top-rung-first wins on this workload.” That label needs unfolding, because either half of it read alone is wrong. The ladder met two conditions: it was cheaper per accepted outcome, and it produced one more correct acceptance. It failed the third: it produced one accepted-but-wrong outcome where top-first produced none. The rule required no increase in wrong acceptances, so the ladder was not adopted and top-first remains the default for this workload under this rule. Not “the strong model is better.” Not “cheap-first lost.” One clause fired, and that clause was written to protect exactly what it protected. 8
The cost column rests on a scenario, not a wage. Executed with the frozen figures:
cost = model_spend + people_asked * PERSON_COST # PERSON_COST = $5, a scenario
cost_per_accepted = cost / accepted
# ladder: (0.0056 + 7 * 5) / 33 = 1.061
# top-first: (0.0317 + 9 * 5) / 31 = 1.453
With five dollars per person, human escalations dominate both totals. The preserved sensitivity runs at $0 (per-accepted costs of fractions of a cent) and $25 ($5.30 against $7.26); the decision is the same at both, and it is also unchanged when top-first’s unknown-spend call is priced in. The ordering is stable across those scenarios; the absolute numbers are scenario artifacts. 8
A05, and the accident that avoided it
The input reads: in staging the cache TTL is one minute; in production it is ten minutes. There is no single correct answer, and the gold says so. The free rung answered 600 seconds, quoting “10 minutes.” Valid JSON, verbatim quote, exactly one duration, seconds matching the quote: the checker passed it, and the acceptance was wrong.
This is Chapter 21’s adequacy lesson in a new setting. The checker verifies grounding. It was never built to verify that the question has one answer, and it did exactly what it was built to do. The checker is not tuned afterward. A05 is evidence of the checker’s limit, and repairing it post hoc would erase the measurement. 8
Top-first did not make the same mistake, and the preserved bytes show why. Its strong-model call finished for length with 2,047 of its 2,048 completion tokens spent on reasoning and empty content. The reasoning is preserved: the model notices the two durations, asks itself “Maybe answer should be null because not a single?”, circles the schema, and is cut off partway through drafting a null answer. The adapter classified the call as empty_output, the call failed, and the item went to a person.
So top-first’s zero came from a token limit reached mid-deliberation, not from a judgment it delivered. The reasoning suggests the model was heading toward null. It never produced an answer the checker could see, and a slightly larger token budget could have produced either outcome. What the run shows is the mechanism of this one avoidance, not better judgment by the strong model. 8
The checker has a mirror-image limit. T07 (gold 7,200 seconds) and T10 (gold 5,400) had correct answers the quote checker could not verify: zero duration expressions found in one quote, two in the other, on every model rung in both arms. Both items went to people after the paid rungs failed them too. A wrong answer passed; correct answers were rejected. What one deterministic checker can verify is not the same as what is true. 8
The rungs that earned nothing
Resolution in the ladder arm: rule 10 of 10 attempts, free model 23 of 30, cheap model 0 of 7, strong model 0 of 7, a person for the remaining 7. The arithmetic closes on its own: 10 + 23 + 7 = 40 items, and 30 + 7 + 7 = 44 model calls, which corroborates the per-item rows independently of the report’s prose. 8
The measured statement stops there. It does not show that stronger models cannot solve such cases, only that on this seven-item remainder the two paid rungs produced no accepted resolution. The recorded reasons suggest why: quote_has_0_durations, quote_has_2_durations, and items with no value to verify — reasons rooted in the input and the checker, which a stronger model cannot change. That motivates a decline-reason-aware policy, routing by why a rung declined rather than escalating every decline through more capacity. It is a hypothesis with a clear next experiment, not a result. 8
The rule rung carries the chapter’s plainest systems point. On the ten structured items, the rule resolved all ten with no model call. Top-first’s strong model got nine right; on S10 it answered with the bare digits 86400 without their key, so the quote contained no duration expression and the item went to a person. That is one item against one model, and the claim stays that size: when the operation is already mechanically specified, a model call adds a way to fail that the rule does not have, without adding anything the rule was missing. 8
Cost is not the only axis. The ladder spent less on models across more calls, took more summed model time (417 s against 315 s, dominated by the free rung’s reasoning), escalated far more often between rungs (51 transitions against 9), and asked people less often (7 against 9). Reducing model spend did not reduce process time. And the scope fence holds on every sentence above: one synthetic workload, one checker, one ladder configuration, one model set, one pricing period, one run per arm, no variance estimate. 8
The spend the ledger could not see
The A05 top-first call returned HTTP 200 with 92 input and 2,048 output tokens reported, and no text. CodeAI’s adapter raised on the missing text, and its failure path then recorded usage as unavailable and spend as unknown, discarding the usage it had been handed. The projection honestly carried one unknown-spend attempt. The independent verifier, recomputing from the preserved bytes at declared prices, found $0.0083 — the most expensive single call in the run — and failed the spend claim. The bundle keeps both: decision tables use known spend with the unknown call flagged beside them. 8
The cause was an ordering bug inside one function: usage was parsed after text extraction, so an extraction failure skipped it. The repair, made while this chapter was drafted and now in CodeAI, parses usage first and carries already-parsed tokens into the failure result. Reduced from the source, with the failure path’s result fields shown as assignments:
raw_usage = parsed.get("usage") if isinstance(parsed.get("usage"), dict) else None
in_tokens, in_reported = _extract_usage(raw_usage, "prompt_tokens")
out_tokens, out_reported = _extract_usage(raw_usage, "completion_tokens")
text = _chat_text(parsed) # may raise; usage is already in hand
...
# failure path
usage_source = "measured" if (in_reported or out_reported) else "unavailable"
cost_usd = estimate_cost_usd(self.model, in_tokens, out_tokens)
cost_source = "estimated" if cost_usd is not None else "unknown"
Three regression tests pin the boundary. The preserved A05 shape (empty content, 92/2,048 usage, a model absent from the pricing table) keeps measured tokens with unknown cost — the $0.0083 belongs to the experiment’s declared prices, not to the adapter’s table. A priced-model variant estimates cost on the failure path. A failure with no usage stays unavailable. The historical bundle is untouched and the experiment was not rerun, so the tables above still show what the run recorded. The sibling adapters share the old ordering and are named, not silently repaired. 11
One adjacent defect stays named. estimate_cost_usd matches model IDs by prefix in table order, so gpt-4o-mini is billed at gpt-4o rates, about seventeen times the intended input price. The one-line fix changes billing semantics under a pricing version that existing tests and the router design pin, so it needs its own versioned migration rather than a quiet edit here. 12 13
The challenger that has not run
The book’s second standing bet is that the process’s routing decision should contain no model call. That bet now has a serious opponent. A deployed system routes in real time between fast and reasoning models using conversation type, complexity, tool needs, and explicit intent, and its router is continuously trained on signals such as users switching models, preference rates, and measured correctness (OpenAI, 2025). That establishes learned routing as a credible production architecture for choosing which model. It does not settle whether operation selection should contain a model call, which is the narrower bet here.
CodeAI’s answer is an experiment design, not a result. It asks whether a model-based router chooses the next operation better than the deterministic policy when both see the same explicit state. Its falsifier says the deterministic bet must be revised if the model-based router is at least 15 percentage points better at operation selection overall, with no stratum worse by more than 5 points; causes no more catastrophic misroutes and no new catastrophe classes; keeps median decision flip rate at or below 10% (p95 at or below 20%); refuses at most 5% of decisions; and produces reasons at least 90% non-vacuous, with auditability not materially worse.
The binding runs the other way too: a deterministic router that is systematically worse is a real failure, because reproducibility alone wins nothing. A second, exploratory track separates where intelligence might pay — interpreting state versus choosing policy — and does not feed the falsifier. The design even states its own power limit: with 24 structured and 24–32 narrative cases, a small correctness difference would be suggestive at best. 14
Its status needs to be classified precisely, because several different things are true at once:
| Artifact | Status |
|---|---|
| Router contract, model-based router, runner, analysis, verifier | Implemented, with tests |
| 68-case corpus | Committed as DRAFT; not frozen; no oracle file |
| Preregistration | Committed; declares itself “NOT FROZEN — NOT RUN” |
| Live comparison | None, under any protocol |
So there is no router accuracy, no catastrophe count, no cost, no flip rate — nothing to report except the design and its falsifier. That is a stronger position than it sounds: the book’s architectural preference is on the table with the exact conditions under which it falls, and nothing in this chapter or the next depends on how the experiment would come out.
Two naming rules hold throughout. The model arm is model-based, never “learned”: nothing was trained on routing data. And in the exploratory track, a deterministic state compiler must answer UNKNOWN where it cannot extract a fact, never guess. Deterministic is not omniscient, and the experiment compares honest boundaries. 14
What this is not
- Not a ladder endorsement. The favored architecture failed its own adoption rule. Cheaper did not mean adopted.
- Not a top-first endorsement. Its clean wrong-acceptance column came, on the one decisive item, from a truncated call.
- Not a model ranking. Rungs are mechanisms with costs. The rule beat the strong model on one structured item because the operation was already specified.
- Not economics. Person costs are scenarios; model prices are one period’s declarations.
Where it is still weak
- One run, one workload, one checker. Forty synthetic items with no repetition; nothing travels without a new run. 8
- The adequacy gap is structural. A05 passes and T07/T10 fail by the checker’s design. 8
- Decline-reason routing is untested. The zero paid-rung resolutions motivate it; nothing measures it.
- The policy ladder is judgment, not a measurement. Its precedence order has never been tested against outcomes;
epistemic-v3is a label, not a validation. 4 - No matrix covers
epistemic-v3. The split-budget, completion and unresolved-effect rows are source-inspected and unit-tested, not enumerated as the v1 matrix was. 17 - Governance binds the operation class, not its parameters. A governed CHECK decision permits a check, not a particular target, command or artifact; and an operation is never accidentally governed, but it can be accidentally external, since naming no decision is the default. 5
- The router track is unrun and not frozen. The contract, harness, analysis, verifier, thresholds, worksheets and a zero-cost instrumentation dry run exist, but the corpus still requires semantic review, no oracle file exists, and
experiments/router-prereg.mdstill declares itselfNOT FROZEN — NOT RUN. After corpus review, two distinct adjudicators must label all 68 cases before a freeze can exist; disagreement remains AMBIGUOUS with no tie-breaker. The full preregistered schedule contains 1,580 model calls, while only the load-bearing R1 tranche — 480 model calls under a $0.25 authorization — is in scope once the prerequisites are satisfied. No provider has been contacted under the protocol, so there is still no router result to narrate. 14 18 - Pricing prefix shadow and sibling adapters. Named, not repaired. 19
Do this now
Thirty minutes. Put one model call on trial before spending it.
- Write the next step of a real task as the seven scheduler facts: process complete, process budget exhausted, unresolved effect, verification pending, proposals wanted, human authority needed, and model budget exhausted. Run them through a function like the one above and write down the operation with its reason. Then change one fact and watch the priority reorder.
- Take an escalation path you already run (retry with a bigger model, then ask a person) and record a decline reason on every transition for one day. Count how many escalations carried a reason no stronger model could fix.
- Find one call your accounting could not see: an empty response, a timeout after billing, dropped usage. Decide what its row should have said, and whether your adapter parses usage before or after the step that can fail.
- Write the falsifier for one architectural preference you hold. If you cannot name the thresholds that would make you revise it, you hold a slogan, not a bet.
If you are building with an assistant:
Decide the operation before the model: a completed process stops; then
process-budget exhaustion stops; an unresolved effect asks a person; pending
verification runs before new cognition; proposals become calls only while
model budget remains; a missing human gate asks a person; otherwise stop.
Keep model and process budgets separate. Climb escalation ladders with a
recorded decline reason at every transition, and adopt a ladder policy only
under a rule frozen in advance that covers cost per accepted outcome, correct
acceptances, and wrong acceptances together. Parse usage before anything that
can fail. Present competing architectures with their falsifiers, and leave
unrun experiments unrun.
Failure modes
- Calling the next model by default. A declined rung is evidence about what to try, not an instruction to spend more.
- Reading one clause of the verdict. “Top-first won” and “the ladder was cheaper” are both true and both incomplete.
- Crediting the cutoff. A call truncated mid-deliberation delivered no judgment, whatever its reasoning was leaning toward.
- Tuning the checker after A05. Repairing adequacy after the fact destroys the evidence that measured it.
- Merging the budgets. One generic “budget exhausted” flag cannot express checks surviving spent cognition.
- Narrating the unrun. Design, implementation, and preregistration are not a result, however complete they look.
What this chapter established
- Choose the operation before the model. Call, check, ask a person and stop are different operations with different costs and failure modes. “Try a stronger model” answers only one of them, and only after something has established that generation is what is missing.
- Keep the three decisions apart. Operation, escalation rung and model choice carry separate evidence: a measured escalation result says nothing about whether the operation policy should be deterministic.
- Split the budgets. Model-budget exhaustion must not stop a check the process still owes. Collapse the two into one budget and the system goes quiet exactly when it should verify.
- Record why each rung declined. A ladder’s cost can only be read against what each rung earned, which means the decline reasons have to be durable.
- Cheaper per outcome is not adoption. A rule frozen before the run is what lets a cheaper, mostly better arm still be refused for the one wrong answer it accepted.
What CodeAI established at these three layers. Operation, escalation and selector have different evidence states: the current scheduler is source-inspected and seam-tested, the ladder has a measured Stage 29B run, and the model-based router remains an unfrozen, unrun challenger design. 4 8 14 The pure decide_next_step function still accepts a caller-supplied SchedulerInput and records nothing. The current runtime path is stronger: ProcessState derives its facts from the ledger, decide_next_for_task records the decision with its state snapshot and basis events, and governance can bind a later CALL or CHECK to that recorded decision. None of those mechanisms measures whether the epistemic-v3 precedence is the best policy. 4 5
The measured ladder was cheaper per accepted outcome ($1.061 against $1.453 at the $5 scenario) and produced one more correct acceptance, and it failed its frozen adoption rule because it added one accepted-but-wrong outcome. 8
A05 shows a grounding checker accepting a grounded wrong answer, and top-first avoided it through a truncated, unusable output rather than demonstrated judgment; T07 and T10 bound the checker from the other side. 8 The failure-path usage defect is repaired in current source with regression tests, the bundle untouched, and the pricing-prefix shadow remains named. 10 The model-router challenge is specified with a falsifier and remains unrun. 14
Evidence notes
Independent verification. The ladder’s verifier imports neither CodeAI nor the producer. It restates the rule, the quote parser and the check, then re-derives rung order, verdicts, decline reasons, spend from usage and declared prices, receipts, both controls, scoring and the decision from preserved bytes: 13 of 14 semantic claims pass, the one failure being the spend under-derivation described above. Five seeded corruptions — a flipped verdict, understated spend, a skipped rung, inflated acceptances, a dropped receipt — each produced problems beyond that baseline failure, a fresh process reprojected all four ledgers identically, and a re-run during drafting gave the same 13/14. 8
The controls also behaved. A scripted wrong answer on the free rung (45,000 seconds for “45000 ms”) failed the check and escalated, and the cheap rung resolved it correctly. Under READ authority, the resolved answer’s write was denied with the apply adapter called zero times; under WRITE, it applied once. 8
The scheduler matrix verifier asserts 16/16 code-versus-table parity at the v1 surface. The router track has a verifier waiting for a run that does not exist, which makes every router sentence above checkable in a simple way: each is design or status, never outcome. 6 14
Next
The pieces are now all on the table, and this chapter deliberately does not assemble them: decision records, authority, observed effects, bound verification, replay, blind proposal collection, measured promotion, and an execution policy with recorded reasons. Each was tested against its own fixture, and several survived by being told no. Whether they hold together when one task has to pass through all of them — and which joints the runtime actually enforces — is the last construction question in the book. Chapter 30 then asks what the construction means once the process, and the policy deciding what happens next, is yours.
Continue with One Process, End to End.
References
- Ehud Karpas, Omri Abend, Yonatan Belinkov, Barak Lenz, Opher Lieber, Nir Ratner, Yoav Shoham, Hofit Bata, Yoav Levine, Kevin Leyton-Brown, Dor Muhlgay, Noam Rozen, Erez Schwartz, Gal Shachaf, Shai Shalev-Shwartz, Amnon Shashua, and Moshe Tenenholtz. MRKL Systems: A Modular, Neuro-Symbolic Architecture That Combines Large Language Models, External Knowledge Sources and Discrete Reasoning. arXiv:2205.00445, 2022. Paper.
- Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M. Waleed Kadous, and Ion Stoica. RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665, 2024. Paper.
- OpenAI. GPT-5 System Card. 2025. System card.
Implementation and evidence sources: Stage 29B remains a historical measured run on its preserved CodeAI revision; current-source claims above were checked against the later successor rather than projected backward into that bundle. src/codeai/scheduler.py: Operation, SchedulerInput, SchedulerDecision, decide_next_step, POLICY_VERSION; src/codeai/process_state.py: ProcessState, BudgetFact, project_process_state, state_snapshot; src/codeai/runtime.py: the raw Runtime.decide_next compatibility path plus process_state, decide_next_for_task, and scheduler-decision recording; src/codeai/governance.py: decision-to-operation binding; src/codeai/ladder.py: run_ladder, _candidate, _check_decline, project_ladder_run, render_receipt; src/codeai/providers.py: OpenCodeCognitionAdapter.send, PRICING_TABLE, estimate_cost_usd; src/codeai/router_model.py, router_contract.py, router_experiment.py, router_analysis.py; experiments/router-prereg.md and experiments/W2-4-prereg-addendum.md. Relevant current tests include tests/test_process_state.py, tests/test_decision_binding.py, tests/test_opencode_gateway.py, and tests/test_budget_safety.py. The scheduler excerpt is current source; the precedence fragments were teaching executions, not a measured policy comparison. Ladder traces and outcomes come from experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/; the older v1 scheduler matrix remains under experiments/applied-ai/evidence/scheduler/. The router design remains unfrozen and unrun: the 68-case corpus has no oracle yet, while experiments/W2-4-dryrun-results.json is instrumentation evidence only. No historical evidence was modified. Footnote prefixes: s inspected source, m pinned run, d preserved demo, r design or frozen report as identified above.
Source inspection:
src/codeai/scheduler.py(decide_next_step, POLICY_VERSION). ↩︎Source inspection:
src/codeai/scheduler.py(decide_next_step). ↩︎Source inspection:
src/codeai/scheduler.py(Operation). ↩︎Source and tests:
docs/seams/scheduler-state.md,src/codeai/process_state.py,src/codeai/scheduler.py(epistemic-v3),tests/test_process_state.py, exampleexamples/applied_ai/ch28_scheduler.py. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Measured run:
experiments/W1-R4-decision-execution-binding.mdand its recheck;src/codeai/governance.py, matrix intests/test_decision_binding.py. Chapter 29 reports the composition audit that found this joint unenforced. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Unpinned demonstration:
experiments/applied-ai/evidence/scheduler. ↩︎ ↩︎ ↩︎Source inspection:
src/codeai/scheduler.py(SchedulerInput). ↩︎Measured run:
experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Source inspection:
src/codeai/ladder.py(run_ladder, _candidate, _check_decline). ↩︎ ↩︎Source inspection:
src/codeai/providers.py(OpenCodeCognitionAdapter.send). ↩︎ ↩︎Source inspection:
tests/test_opencode_gateway.py. ↩︎Source inspection:
src/codeai/providers.py(PRICING_TABLE, estimate_cost_usd). ↩︎Source inspection:
tests/test_budget_safety.py. ↩︎Report:
docs/applied-ai/ch28-router-experiment-design.md. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Source inspection:
src/codeai/router_model.py,(router_contract.py). ↩︎Source inspection:
experiments/router-prereg.md. ↩︎Source inspection:
src/codeai/scheduler.py(POLICY_VERSION). ↩︎Preparation, not result:
experiments/W2-4-prereg-addendum.md, oracle worksheetsexperiments/W2-4-oracle-worksheet-adjudicator-{a,b}.json, instrumentationexperiments/W2-4-dryrun-results.json. Criteria unchanged fromdocs/applied-ai/ch28-router-experiment-design.mdandexperiments/router-prereg.md. ↩︎Source inspection:
src/codeai/providers.py(PRICING_TABLE). ↩︎