The Agent Cannot Grade Its Own Homework
Asserted is not executed, executed is not passed, and passed is not established. A check must be independent of the generator's claim, adequate to the property that matters, and bound to the exact state being accepted — and each of those can fail while the others hold.
Part 4 — Make It Safe and Verifiable
Independence, adequacy, and binding
A model saying its work is correct is not verification. Neither, on its own, is running a separate test. A test can run separately and check the wrong thing. It can check the right thing against a version of the artifact that has since changed. And a PASS that nobody can tie to the exact state being accepted is only another claim.
asserted ≠ executed ≠ passed ≠ established
A verification step you can rely on has three properties, and each one can fail while the other two hold:
- Independence. The check is separate from the generator’s own success claim.
- Adequacy. The check tests the property that actually matters.
- Binding. The check ran against the exact state, artifact and claim being accepted.
By the end of this chapter you will be able to examine any AI verification pipeline, your own included, and ask those three questions separately. You will also see what a PASS does not prove, even when all three hold.
The patch that passed
Picture a patch arriving with a note: “all tests pass.” The note is confident and unchanged while you read it. Then you run one small check against the actual file, and it fails.
Keep both in front of you. The sentence says success. The command says failure. One of them examined the artifact; the other one is a report about it. That is independence. Adequacy and binding are still open: the failing check might test the wrong property, or an older version of the file.
Chapter 20 ended with permission: who may invoke the worker. Permission does not establish that the resulting state satisfies the property the task cares about. A worker that was allowed to act can still produce the wrong bytes, and a worker that claims success can still be wrong about what it did.
What observation can establish that an action met its check?
Asserted, executed, passed, established
Each stage answers a different question and leaves another open:
| Stage | What happened | What remains open |
|---|---|---|
| Asserted | Someone, or some model, said the work succeeds | Whether anything examined the artifact |
| Executed | A check ran against a target | What verdict it returned |
| Passed | This verifier returned PASS for this request, state, and environment | Whether it tested the property you care about, on the state you are accepting |
| Established | A bounded property holds for an identified state, within the check’s scope | Everything outside that scope |
Three properties make the last row precise. They are independent axes, not three names for care:
| Property | Question |
|---|---|
| Independence | Is the check separate from the generator’s own success claim? |
| Adequacy | Does the check test the property we actually care about? |
| Binding | Did the check run against the exact state, artifact, and claim now being accepted? |
A verifier can be independent but inadequate: it runs separately and tests the wrong thing. It can be adequate but unbound: it tests the right property against stale state. It can be correctly bound but irrelevant: the right artifact, the wrong question. The rest of the chapter shows each failure with its companions held constant, from a measured experiment and a pinned run.
The evaluation literature earns its place here in exactly this shape. SWE-bench tasks a model with editing a real codebase to resolve real GitHub issues and grades the result with the repository’s own tests: 2,294 problems across 12 Python repositories, with the best early model resolving under 2% (Jimenez et al., 2024). That supports checking changes by execution rather than accepting a completion statement. It does not establish that arbitrary tests prove correctness; the verdict is bounded by the tests’ coverage and their environment. The mapping to CodeAI’s check boundary is the book’s own.
Cobbe and colleagues separate candidate generation from verifier-based selection on grade-school math: they generate many candidate solutions and keep the one a trained verifier ranks highest, and verification improves with data better than the fine-tuning baseline (Cobbe et al., 2021). The relevant lesson is role separation, not infallibility. Their verifier is learned, not a deterministic proof checker; a learned judge ranking candidates is not ground truth about any single answer. The application of that separation to CodeAI’s local checks is the book’s mapping.
Zheng and colleagues measure LLM-as-judge agreement against human preferences and find strong judges matching humans at roughly the human–human rate, alongside position, verbosity, and self-enhancement biases with limited reasoning ability (Zheng et al., 2023). Preference agreement is a boundary case for this chapter: useful where exact checks are unavailable, but agreement about which answer reads better must not be presented as proof that an artifact behaves correctly. That boundary is the book’s use of their result, not theirs.
How CodeAI runs a check
The verification types are small. CheckVerdict has four members: PASS, FAIL, INCONCLUSIVE, and ERROR. CheckRequest carries the check’s identity, its task, the claim IDs it targets, the command with working directory, timeout, and environment overrides, plus target, target_state_hash, and environment_hash. CheckResult carries the verdict with exit code, preserved stdout and stderr, details, error, and — since a revision made for this chapter — the runtime’s pre-check state reading. 1
Current LocalCommandVerifier.run can produce all four verdicts under a mapping declared before execution. An empty command, timeout or operating-system launch failure is ERROR. Otherwise an ExitCodePolicy maps exit codes: by default 0 is PASS and other codes are FAIL, while a check can explicitly declare codes as INCONCLUSIVE or enumerate its FAIL codes. An exit code outside an explicit mapping is ERROR rather than a guess, and INCONCLUSIVE must carry a reason. The subprocess inherits a copy of the process environment with the request’s overrides applied; a short override list is not a hermetic environment. Standard output and standard error are captured and returned on the result. 2
The verifier receives a CheckRequest, but the runtime owns the bindings that say what subject that request actually reached. Current Runtime.run_check delegates to codeai.verification.run_check. The current successor also passes through the later operation-governance seam before verification; once execution is permitted, the Chapter 21 path is reduced from the source like this:
binding = bind_check_target(runtime, request)
with TemporaryDirectory(prefix="codeai-check-") as workdir:
artifact_binding, prepared = bind_check_artifact(runtime, request, workdir)
result = execute_verification(
runtime, prepared, binding, verifier, artifact_binding
)
result = finalize_verification(
runtime, result, binding, artifact_binding
)
completed = record_verification(
runtime, request, binding, result, request_event, verifier, artifact_binding
)
record_verification_attempts(runtime, request, completed)
apply_verification_to_claims(runtime, request, result)
State binding reads the resolver exactly once when target_state_hash is requested. If no state binding was requested, the check is explicitly UNBOUND and no state read is taken. An unavailable, raising or mismatched requested state produces ERROR before verifier invocation. Artifact binding can independently stop the check before invocation when named bytes are missing, corrupt or never supplied through the declared command interface. A raising, malformed or unexplained-INCONCLUSIVE verifier result is likewise normalized to ERROR. Finalization overwrites any verifier-supplied account of the observed target state with the runtime’s own reading. 3 4
check.completed records the result beside the state binding, artifact binding, declared verdict policy and verifier identity. PASS and FAIL can change the named claims under the legacy promotion path; INCONCLUSIVE and ERROR do not. A separate attempt relation is still recorded for every named claim, so “produced no evidence” no longer means “no attempt is discoverable.” 5
The single-read repair matters because the version Stage 29B ran on could call _current_state_hash() multiple times, and a missing resolver could let a requested binding fall through as if none had been requested. The older evidence bundles preserve that historical behavior; the pinned verification run later in this chapter exercises the repaired state-binding path.
The generator’s claim and the verifier path remain separate, while state and artifact binding answer different questions before a verdict is allowed:
The legacy promotion itself is the chapter’s named conflict. On PASS, every claim the request names receives claim.evidence at E3_REPRODUCED; on FAIL, every named claim receives a refuting claim.status. The runtime does not ask whether the command tested anything about those claims. Chapter 18 showed the shape of this failure: a claim promoted by a check whose command printed “nothing tested.” The newer evidence path can require a completed check that named the claim, but naming is still not testing.
The runtime can therefore record support from a targeted PASS without establishing that the checker was adequate to the claim. Reading that record as warranted support still requires the adequacy argument this chapter keeps outside the request schema. This revision leaves that legacy promotion path visible rather than turning a binding repair into a claim-semantics repair. 5
Independence: a check that contradicts the claim
A small illustrative fragment, executed as teaching code with the real APIs: 3
generator_claim = {
"status": "success",
"tests_passed": True,
}
result = runtime.run_check(
CheckRequest("check-1", "task-1", claim_ids=("claim-1",)),
verifier=verifier,
)
assert result.verdict == CheckVerdict.FAIL
assert "real stdout" in (result.stdout or "")
The generator’s sentence is left untouched while the target is examined, so its claim and the verifier’s observation sit side by side, and the contradiction is visible without anyone having to judge the model’s tone. The FAIL verdict with preserved stdout is what independence buys: not correctness, just a second observer that refuses to take the first one’s word.
Adequacy: a grounded answer that was wrong
Stage 29B ran a 40-item TTL corpus through two arms: a ladder of rule, free, cheap, and strong rungs against sending every item to the strong model first. The ladder was cheaper per accepted outcome, but it failed its frozen adoption rule because it accepted one wrong answer where the other arm accepted none. That wrong answer is this chapter’s business. 6
The in-process check used by the execution ladder follows the same verdict discipline: the caller’s deterministic check runs, its boolean becomes PASS or FAIL, and a crashing check becomes ERROR, never PASS. 7
The Stage 29B TTL checker, ttl-check-v1, is deliberately narrow and stays that way for this chapter. It requires the output to parse as one JSON object holding an integer seconds and a non-empty quote; the quote must appear verbatim in the input; the quote must contain exactly one duration; and the seconds must equal that duration. Anything else fails with a named reason: output_not_json, no_value_to_verify, seconds_not_a_whole_number, quote_missing, quote_not_in_input, quote_has_0_durations, quote_has_2_durations, or seconds_do_not_follow_from_quote. Those semantics are frozen historical evidence about that stage, not a component this chapter tunes. 8
A05 is the chapter’s central adequacy case, and the checker is left exactly as it was. The input states two TTLs: 1 minute for staging, 10 minutes for production, with no single correct answer. The free rung answered {"seconds": 600, "quote": "10 minutes"}. The quote appears verbatim in the input, it contains exactly one duration, and 600 follows from it — so the check passed, and the answer was accepted and wrong against the frozen gold.
Grounding and derivability were verified; ambiguity and relevance were not. The check answered “does this value follow from a quoted passage?” while acceptance needed “does this input determine one value?” Those are different properties, and no tuning of the duration parser would close the gap without changing what the check claims to test. 9
T07 and T10 are the complementary false rejections, also preserved untuned. T07 reads “Cache TTL: 7200 (seconds). Connection pool idle timeout: 600 seconds.” with gold 7200; the parser finds no duration expression in a quote of the form 7200 (seconds), so correct answers were declined as quote_has_0_durations. T10 reads “Entries are cached for 1 hour 30 minutes” with gold 5400; the quote holds two durations, declined as quote_has_2_durations.
Both items went to a person after repeated model calls in the ladder arm. Correct answer ≠ answer verifiable by this checker. The preregistration had already recorded both as properties of the check, kept rather than tuned away. 9
The forced-wrong control shows the check doing one part of its declared job: a scripted 45000-seconds answer failed seconds_do_not_follow_from_quote, the run escalated, and the next rung resolved the item correctly. That seeded defect is a useful negative control because it shows sensitivity to at least that failure class. It does not, by itself, establish that the checker is adequate to the whole task. 9
Together the three cases separate the axes. A05 was independent and bound but inadequate. T07 and T10 were independent and bound with an inadequate parser in the other direction — strict where the property needed leniency. The forced-wrong case was adequate to its narrow property and correctly bound. None of these sentences can be shortened to “the verifier worked” or “the verifier failed” without losing which property held.
Binding: refusing to check the wrong state
Each fragment below was executed as teaching code with the real APIs, and the assertions held: 3
Requested state state-B, observed state-A: the verifier never runs.
result = runtime.run_check(
CheckRequest("check-2", "task-1", target_state_hash="state-B"),
verifier=verifier,
)
assert result.verdict == CheckVerdict.ERROR
assert verifier.calls == 0
A request that requires binding with no resolver configured returns ERROR, with no observation invented and no verifier call.
result = Runtime(SQLiteLedger()).run_check(
CheckRequest("check-3", "task-1", target_state_hash="state-A"),
verifier=verifier,
)
assert result.verdict == CheckVerdict.ERROR
assert result.error == "target state unavailable: requested binding cannot be checked"
assert verifier.calls == 0
Matching binding lets the verifier run, and the reading that permitted it stays on the result:
result = runtime.run_check(
CheckRequest("check-4", "task-1", target_state_hash="state-A"),
verifier=verifier,
)
assert result.observed_target_state_hash == "state-A"
The regression tests behind this change cover the further cases: a resolver that raises (observation failure, ERROR, zero calls), a resolver read exactly once even when the state moves between reads, an unbound check that takes no reading at all, a verifier-supplied observation overwritten by the runtime’s own, and a verifier that raises (durable ERROR with the exception preserved, exactly one invocation, completion surviving ledger reopen, no claim promoted). 10
Two bindings, not one
State is one subject a check can have. Bytes are another, and for a while they were held to a weaker standard. A check said which artifact it was about by carrying a label — artifact:sha256:… — and acceptance compared that label with the artifact being accepted. A label checked against a label. A command that was sys.exit(0), which never opened the artifact, could carry its name, pass, and complete an acceptance. 11
The artifact is now a reading too. The runtime resolves the named bytes from the store, verifies their digest, writes them where the check can read them, and substitutes {artifact} in the command with that path, recording the result beside the state binding:
| Artifact binding | Meaning | Verifier runs? |
|---|---|---|
BOUND | Resolved, digest verified, supplied to the check | yes |
UNBOUND | The check names no artifact | yes |
MISSING | Named, and not retrievable | no — ERROR |
MISMATCH | The stored bytes do not hash to the named digest | no — ERROR |
UNCONSUMED | A command check that never references {artifact} | no — ERROR |
UNCONSUMED is the one worth sitting with. The artifact existed, the digest verified, and the command still never received it through the interface it declared. There is no verdict to report, because the measurement never had its subject: a malformed measurement is an ERROR, not a result. That is the same rule as a binding failure, applied to bytes instead of state. 12
And it stops precisely there:
artifact binding ≠ artifact adequacy
BOUND means the intended artifact was resolved, integrity-checked and made available through the declared interface. It does not mean a single byte was read. A command that greps the artifact for a required token and one that takes the path and ignores it are both legitimately BOUND, and only the first is a check of anything. Establishing what a verifier actually read would mean sandboxing or tracing its execution — a different problem, and not one this boundary pretends to solve. 12
An attempt is not evidence
A check that errors produces no evidence about the claim it named — correctly, since verification failed ≠ claim disproved. For a long time it also produced no record against that claim, which made two different facts identical:
no evidence was produced ≠ no attempt is discoverable
Those are now separate. Every verification aimed at a claim is recorded against it, whatever the outcome, as a reference to the check rather than a copy of its verdict — so the claim side can never come to disagree with the check it describes. The hostile case runs end to end:
claim recorded
-> a check names it, and names an artifact that was never stored
-> artifact binding MISSING
-> the verifier never runs
-> ERROR
the claim's attempt history contains that check, with its verdict read from it
the claim's evidentiary state unchanged
What is enforced here is narrow and worth naming exactly: the linkage and its admissibility rules — an unknown claim is refused rather than linked, one check yields one attempt, and ERROR and INCONCLUSIVE are inadmissible as evidence. The claim’s own evidentiary state remains derived from the recorded checks. The runtime does not establish that any evidence is true. 13
All three properties in one pinned run
The pinned run for this chapter used a protocol frozen before execution, real local commands against a disposable fixture, and no network. A separate stdlib-only verifier reconstructs every row — requested binding, recorded observation, verdict, invocation count, and claim effects — from the ledger, the fixture bytes, and the receipts. The success criterion was not “all PASS”: the frozen expectations require the appropriate PASS, FAIL or ERROR for each constructed case, and the verifier checks those recorded relationships rather than treating PASS as the desired universal outcome.
| Case | Epistemic state |
|---|---|
| Genuine claim contradicted by an independent check | asserted claim.recorded sits beside a FAIL; the ledger then carries claim.status REFUTED — said and established stay separable |
| Weak vs strong oracle, same artifact | presence check PASSes (with claim.evidence REPRODUCED) while the exact-bytes check FAILs (with REFUTED after REPRODUCED, in ledger order) |
| Stale binding | ERROR naming the mismatch, the observation preserved as the mutated-bytes hash, zero verifier invocations |
| Unavailable resolver, two ways | ERROR saying unavailable (None) and ERROR naming RuntimeError (raising), both with zero invocations |
| Raising verifier | durable ERROR with the verifier raised prefix and the type, distinct from a target failure — and the one invocation that produced it |
| Procedure failures | 30-second sleep on a 1-second budget times out; a nonexistent binary fails at spawn with no exit code — both ERROR, never PASS |
| Negative controls | the exact-bytes checker PASSes on correct bytes and FAILs on the seeded defect before its PASS counts |
| INCONCLUSIVE | absent from all ten completions: the verifier version exercised in this pinned run produced PASS/FAIL by exit code and ERROR otherwise, so the stage recorded the gap instead of fabricating a case; current LocalCommandVerifier later gained declared INCONCLUSIVE mappings |
The run’s limits stay with its result: adequacy unenforced, no hermetic execution, observation not a lock, and concurrent mutation between observation and verification still belongs to retry policy.
What this is not
- Not correctness. A PASS is this verifier’s verdict on this request, state, and environment — never a proof of the property the reader cares about.
- Oracle design stays elsewhere. The runtime can refuse unbound checks; it cannot tell whether a command tests the right property. Adequacy stays with whoever wrote the checker.
- Not hermetic execution. The environment is inherited plus overrides, and the state reading is an observation, not a lock.
- Naming is not support. A passing check that names a claim promotes it under the legacy path whether or not the command tested it.
- A verdict is not an oracle.
INCONCLUSIVEnow has a real producer, but a check declaring that it cannot settle a question is still only that check’s report of its own limits.
Where it is still weak
- Adequacy is unenforceable at the boundary. Nothing in the request certifies that the command tests the claim it names. 5
- Verdict mappings are declared, not semantically certified.
INCONCLUSIVEhas a producer now: a command verifier maps exit codes under a policy declared before execution, the request and result preserve the mapping identity, and an unmapped code is ERROR rather than a guess. The runtime can check that the reported policy is consistent with the declared one; it cannot establish that mapping exit code 2, or any other code, to INCONCLUSIVE is the right interpretation for that command. That is adequacy again. 2 - The environment is not captured. Overrides are recorded on the request; the inherited remainder is not. 2
- A hash identifies; it does not freeze. Concurrent mutation between observation and verification, or between verification and acceptance, is outside the binding record. Retry policy owns this interval. 3
- The resolver is trusted configuration. A constant or partial resolver supplies a poor reading, and the runtime does not validate its coverage. 14
- Binding is supply, never consumption. The runtime can establish which bytes and which state a check was given; nothing short of instrumenting the verifier establishes what it read. Adequacy enforcement, hermetic environments and mutation-proof binding across the observation–verification interval all remain open. 12
Do this now
Thirty minutes. Take one success claim and refuse to believe it.
- Pick a generated artifact with a confident completion sentence. Keep the sentence unchanged and write a check that examines the artifact’s bytes, not the sentence.
- Run the check. Preserve the command, working directory, environment overrides, exit code, stdout, and stderr beside the claim.
- Weaken the checker deliberately until it passes something wrong — syntax for semantics, presence for correctness. Write down which property each version tests.
- Bind the check: request a target state hash, then change the target and re-run. Confirm the stale request refuses before the command runs, and note what the refusal does and does not prove.
If you are building with an assistant:
Separate the generator's success claim from verification.
Run an independent check against the artifact's bytes and preserve
command, working directory, environment overrides, exit code, stdout,
and stderr. Test the checker against seeded defects before trusting its
PASS. Bind each check to an exact target state with a single runtime
observation; refuse stale or unreadable states before invoking the
verifier, and record the observation used. Do not present PASS as proof
of adequacy, and do not describe a later passing check as caused by a
repair unless the same check observed the repaired target.
Failure modes
- Grading its own homework. Accepting the generator’s success sentence as the verification.
- Reading PASS as proof. Treating one verifier’s verdict as establishment of the underlying property.
- Testing the wrong property well. A deterministic, grounded, perfectly implemented check for something nobody needed.
- Checking yesterday’s artifact. Running a strong oracle against state that has since moved.
- Counting the repair twice. Describing a later different check’s PASS as caused by a fix the check never observed.
- Promoting by naming. Letting a passing check support claims it never tested because the request listed them.
- Confusing the verifier’s failure with the artifact’s. An ERROR says the checking procedure did not complete; it is not a FAIL, and it still belongs in the ledger.
- Hiding procedure failure in ERROR. Treating “the checker crashed” as neutral when acceptance still needs an answer.
What this chapter established
- Asserted ≠ executed ≠ passed ≠ established. A PASS is one verifier’s verdict under one declared check and environment. If state or artifact binding was requested, the record must also establish what subject the verifier was given. None of that proves properties outside the check’s scope.
- Independence. The generator’s success claim is not verification. A separate check can contradict it, but separation alone says nothing about whether the check is adequate.
- Adequacy. A check can be separate, correctly bound, and still test the wrong property: grounding instead of ambiguity, syntax instead of meaning, presence instead of exact bytes. Nothing on a request certifies adequacy. A seeded defect can demonstrate sensitivity to one known failure class; it cannot establish adequacy by itself.
- Binding. When state binding is requested, take one reading, refuse unavailable or mismatched state before invoking the verifier, and preserve the reading used. When artifact binding is requested, resolve and integrity-check the named bytes and supply them through the declared interface. Supply is still not proof of consumption, and neither binding freezes concurrent state.
- A broken checker is not a failed artifact. A verifier that crashes, times out, returns a malformed result or produces an unmapped outcome yields ERROR. A legitimately executed check that cannot settle the question may instead return INCONCLUSIVE under a declared policy; neither outcome is FAIL.
What CodeAI showed. In Stage 29B, the answer to A05 was checked by a deterministic checker separate from answer generation, but the checker was inadequate to the task: it verified that 600 seconds followed from a quoted passage while the input stated two TTLs, and a wrong answer was accepted. The same strict checker rejected correct answers to T07 and T10, and a seeded wrong answer was caught; the checker’s semantics were frozen, not tuned. 9
The pinned ten-check verification run exercised the repaired state-binding path with PASS, FAIL and ERROR outcomes expected by its frozen protocol; a separate verifier rejected five targeted seeded corruptions. That run contained no INCONCLUSIVE completion. Current CodeAI goes further: LocalCommandVerifier can produce a reasoned INCONCLUSIVE under a declared exit-code policy, artifact binding distinguishes supplied bytes from a caller-written label, and every named claim can retain a verification-attempt trace even when no admissible evidence was produced. Adequacy is still not enforced, execution is not hermetic, and binding does not lock state against concurrent change.
Evidence notes
The pinned run. The verification-binding bundle is experiments/applied-ai/evidence/verification-binding/2026-09-14-1b3c7a2/. Seeded corruptions — a flipped verdict, an altered observation, a dropped claim effect, an ERROR rewritten to PASS, an inflated receipt — are each rejected with the failure named.
The earlier verifier-independence demo. The preserved verifier-independence demo runs six local checks with no network and no models, and its verifier wrapper counts invocations. 15
| Case | What ran | Verdict | What it establishes |
|---|---|---|---|
| A | Generator JSON claims success; the check command exits 1 | FAIL, real stdout preserved | Assertion ≠ observation |
| B | A later, different check command exits 0 | PASS | A later check passed; see the caveat below |
| C | A JSON-syntax check on semantically wrong {"threshold": 9999} | PASS | The oracle certifies syntax only |
| D | Sleep-30 with a 0.5 s timeout; a missing binary | ERROR / ERROR | Procedure failure, not a verdict on the claim |
| E | Requested state-B against resolver state-A | ERROR, 0 added verifier invocations | Stale targets never reach the command |
Case B needs its limitation stated plainly. The failing and passing commands were different commands, and neither read or bound itself to the patch.txt file whose “repair” the narrative discussed. The bundle preserves one constructed check failing and another, later check passing; it does not establish that a target repair caused the change. The repair happened and the later check changed, but the same check observing the repaired target and passing is not in the evidence. The teaching point survives the narrowing: passing results do not travel across revisions, and a state hash identifies a target without making concurrent mutation impossible. 15
Case C is the adequacy exhibit in miniature: a syntactically valid but semantically wrong artifact passes a syntax-only checker. Independence held — the checker ran separately from whatever produced the file — while adequacy failed, and the PASS certifies JSON syntax and nothing about the threshold. 15
Independent verification. The demo’s independent verifier imports no CodeAI code for its predicates. It reads the preserved summary and requires A to FAIL with stdout, B to PASS, C to PASS as a bounded oracle, D to be ERROR/ERROR, and E to short-circuit with zero added invocations; it also mutates a copy with a flipped A verdict and requires rejection. 16
That independence has a limit: the verifier reads summary fields without re-running the commands or re-reading the targets, so it would catch a mislabeled verdict but not a command that never examined its target — the case-B limitation above, found by reading the producer script rather than by any predicate. An independent check of a summary is not an independent repetition of the work. 15
Stage 29B’s own verifier checked harder. Five seeded corruptions — a flipped check verdict, zeroed spend, a removed rung, inflated acceptances, a dropped receipt — were each rejected with new problems beyond the baseline failure, and fresh reprojection of all four ledgers matched. The flipped-verdict corruption is the one this chapter leans on: T07’s recorded PASS re-checked as quote_has_0_durations, so a tampered verdict fails re-derivation. 9
Next
The process can now refuse to check the wrong state, and it knows what its checks do not prove. Refused and failed work will come back for another attempt — and every retry re-enters the world through effects that must be controlled.
Continue with Retries Are Side Effects Too.
References
- Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR, 2024. arXiv:2310.06770.
- Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168, 2021. Paper.
- Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks, 2023. arXiv:2306.05685.
Implementation and evidence sources: historical behavior and Stage 29B remain frozen to the versions that produced those records; current-source statements above were checked against the later CodeAI successor. src/codeai/adapters.py: CheckRequest, CheckResult, CheckVerdict; src/codeai/verifier.py: current LocalCommandVerifier and its declared exit-code policy behavior; src/codeai/verification.py: state binding, artifact binding, ExitCodePolicy, verifier-result validation, recording, claim promotion and verification-attempt tracing; src/codeai/runtime.py: the run_check façade and _current_state_hash; src/codeai/ladder.py: historical _InProcessCheck; src/codeai/evidence.py: Chapter 18 claim-evidence discipline; src/codeai/governance.py: the later operation-governance seam that now precedes verification in current source. Teaching fragments were illustrative working-tree executions, not pinned stage runs. Evidence remains experiments/applied-ai/evidence/verifier-independence/, experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/, and experiments/applied-ai/evidence/verification-binding/2026-09-14-1b3c7a2/, with later artifact-binding and claim-attempt work named in the footnotes. Footnote prefixes: s inspected source, m pinned run, d preserved demo, r frozen report. Later source behavior is not retroactively attributed to the earlier evidence bundles, and the Stage 29B TTL checker semantics remain unchanged.
Inspected code:
src/codeai/adapters.py(CheckVerdict, CheckRequest, CheckResult). ↩︎Inspected code:
src/codeai/verifier.py(LocalCommandVerifier.run). ↩︎ ↩︎ ↩︎Inspected code:
src/codeai/runtime.py(Runtime.run_check). ↩︎ ↩︎ ↩︎ ↩︎Inspected code:
src/codeai/runtime.py(Runtime.run_check, Runtime._finalize_check_result). ↩︎Inspected code:
src/codeai/runtime.py(Runtime._apply_check_to_claims). ↩︎ ↩︎ ↩︎Report:
experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/chapter-evidence-report.md. ↩︎Inspected code:
src/codeai/ladder.py(_InProcessCheck). ↩︎Report:
experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/run.py. ↩︎Measured run:
experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Binding tests:
tests/test_check_binding.py. ↩︎Measured run: CodeAI’s frozen Wave 1 composition audit, probes I and M (
experiments/W1-composition-results.md, baseline runtimef7d2910). Chapter 29 reports the audit in full. ↩︎Measured runs:
experiments/W1-R3-artifact-binding.mdand the verification seam it extends;src/codeai/verification.py(ArtifactBinding,ExitCodePolicy), limits indocs/seams/verification.md. ↩︎ ↩︎ ↩︎Measured run:
experiments/W1-R5-claim-attempt-trace.md;src/codeai/verification.py(verification_attempts_for_claim). ↩︎Inspected code:
src/codeai/runtime.py(Runtime._current_state_hash). ↩︎Unpinned demonstration:
experiments/applied-ai/evidence/verifier-independence. ↩︎ ↩︎ ↩︎ ↩︎Unpinned demonstration:
experiments/applied-ai/evidence/verifier-independence/verify_verifier.py. ↩︎