← Applied AI

The Agent Said Done. Did Anything Change?

Decided is not requested, requested is not performed, and a worker's report of what it performed is not an observation of what changed. Keep the report and a separately sourced reading of resulting state apart; a success report and a changed-state reading are different facts, and neither alone establishes that the intended change happened. Tested with honest, lying, partial and wrong-target workers.

Part 3 — Give Intelligence a Runtime

Observe the effect, not the report

Sooner or later an AI application is allowed to change something: edit a file, update a record, send a message, merge a change. The moment it does, a new kind of mistake becomes possible. A worker reports “done”, the report is recorded as the outcome, and nobody looks at what actually changed. The report can be sincere and still wrong. The edit went to the wrong file, stopped halfway, or never happened.

decided  ≠  requested  ≠  performed  ≠  observed

A worker’s “done” is another claim, the same kind of thing Chapter 18 refused to treat as evidence. A separate reading of resulting state can corroborate or contradict that claim within the reader’s declared scope. To remain separate evidence here, it must not simply reuse a state value supplied by the worker. Even then, observing a change is not the same as establishing that the intended change occurred.

By the end of this chapter you will be able to make an AI action inspectable from request to effect: record the request before anything runs, keep the worker’s report apart from a separately sourced observation of the result, compare that observation with what was intended, and recognize the honest, lying, partial and wrong-target workers that a success status cannot tell apart.

The edit that has not happened

Picture a small case. A reviewer proposes changing a cache setting. The patch is in an artifact. A person decides to apply it. The file still contains the old value.

Then a worker returns “done” and nobody opens the file.

The decision might have been well supported, the instruction exact, and the worker’s return normal — none of those tells the person reading the record whether the file changed.

Chapter 18 ended: “A decision can now say what it rested on. It still changes nothing.” This chapter crosses that boundary, but it needs to cross it in pieces.

How does a proposed action become an observed effect?

Decided, requested, performed, observed

Each stage answers a different question and leaves another open:

StageQuestion answeredWhat remains open
DecidedWhat change was chosen, and on what basis?Has anyone asked a worker to do it?
RequestedWhat operation was submitted, by whom, against which state?Did execution begin?
PerformedWhat did the worker do?What evidence of that effect survives?
ObservedWhat did a specified reader find afterwards?Does that observation establish the intended result?

“Performed” is a fact about an effect. A worker’s statement that it performed something is a report about that fact. Keeping the report is useful; promoting it silently into independent evidence is the mistake.

An effect report is another claim. Observation is evidence about the effect. This carries Chapter 18’s distinction across the execution boundary: what the actor says and what an observer reads must remain separately inspectable.

That gives a ladder with four rungs rather than two, and each rung is a different fact:

the worker says done  ≠  something changed  ≠  the intended thing changed
                      ≠  the intended result was established

CodeAI keeps these as separate dimensions across the recorded result and the action-state projection, plus verification’s answer:

DimensionValuesWhose answer it is
Result statusSUCCEEDED / FAILED / DENIEDthe actor’s report
State observationCHANGED / UNCHANGED / UNAVAILABLEa projection over the runtime’s before and after resolver readings
Effect stateOBSERVED / UNKNOWN / REPORTED / NONEthe conservative conclusion the action record supports
VerificationPASS / FAIL / INCONCLUSIVE / ERRORa declared check (Chapter 21)

Effect state is derived from result status and the available readings, conservatively. Reported success with a changed scope is OBSERVED: the selected representation changed across the effect boundary, not necessarily in the intended way. Reported success with an unchanged scope is UNKNOWN rather than “nothing happened” — the effect may have landed outside what the resolver watches, or produced the same representation — and it asks for reconciliation. Reported success with no observation available is REPORTED: the actor’s claim, carrying neither corroboration nor contradiction. 1

ReAct interleaves model reasoning with actions that obtain information from external sources and environments. That supports treating an action and the information returned from acting as distinct steps. Its benchmark results do not establish durable recording or authorization for this runtime. The mapping to CodeAI’s action boundary is the book’s own (Yao et al., 2023).

PAL makes a related separation: a language model generates intermediate programs and an interpreter executes them. Moving computation into an interpreter changes who performs it. It does not, by itself, establish that a proposed file edit meets a user’s objective. That application to effectful software is also the book’s mapping, not a result reported by PAL (Gao et al., 2023).

The construction here needs ordinary software around that boundary: a request, a callable worker, a record of its answer, and a reading whose origin is explicit.

Write down the request

CodeAI’s ActionRequest separates the operation’s identity from its task, instruction, capability, precondition, execution attribution, and optional evidentiary justification. directive_id can associate it with a directive; payload carries application data; decision_id can name the Chapter 18 decision whose recorded basis this action claims to rely on. idempotency_key is also present, but its consequences belong to Chapter 22. 2

The pinned Chapter 19 evidence used an earlier application convention in which a decision reference lived in payload. The current type has a dedicated field. Keep those two evidence states separate rather than reading today’s source backward into the historical run. Here is the current shape; the identifiers and proposed value are illustrative, and creating the object executes nothing. This fragment assumes before_hash was computed from the intended target under a declared policy. 2

Where the effect happens

The execution interface is small. Here is its signature, with the surrounding imports omitted: 3

class ExecutionAdapter(Protocol):
    def execute(self, request: ActionRequest) -> ActionResult: ...

The adapter is where the side effect belongs. The runtime supplies an explicit request and receives an ActionResult. That result can carry a status, timestamps, transcript, stdio, exit code, and state-related fields. ActionStatus has SUCCEEDED, FAILED, and DENIED; there is no separate “objective verified” member. 4

The pinned Chapter 19 run captured an earlier form of this action boundary. Current execute_action has since acquired the authority, recovery, decision-evidence and governance seams developed later in the book. The current order is worth naming without teaching those later chapters early: 5

OrderOperationMeaning at this boundary
FirstAppend action.requestedPreserve what was submitted
Pre-effect gatesRecord governance source, resolve authority, and validate a cited decision basis when presentLater chapters own these policies; all run before the adapter
Reuse checkLook up a completed prior result by keyMay return a reused result; Chapter 22 examines the conditions
New executionRead _current_state_hash() and compare a supplied preconditionCompare the planning state with an available current reading
Commit boundaryAppend action.execution_started, including the runtime’s before-state readingAfter this point an interrupted record may represent an effect of unknown outcome
EffectCall adapter.execute(request)Hand control to the worker
After returnRead _current_state_hash() again; finalize the resultKeep the worker’s report apart from the runtime reading
LastAppend action.completedPreserve the completed result record

The boundary as a flow — the report and the reading travel separate paths:

Read the file after the worker

Here is the whole boundary in one disposable fixture, reduced and using the real APIs: a request, a worker that writes, and a read of the target that the worker does not control. It is teaching code rather than the pinned run reported next. The worker deliberately returns a report string where a digest might be expected; the runtime keeps its own observation in a separate field. 6

from hashlib import sha256
from pathlib import Path
from tempfile import TemporaryDirectory

from codeai.adapters import ActionRequest, ActionResult, ActionStatus
from codeai.domain import Authority, Capability
from codeai.ledger import SQLiteLedger
from codeai.runtime import Runtime

with TemporaryDirectory() as directory:
    target = Path(directory) / "cache.toml"
    target.write_bytes(b"ttl_seconds = 60\n")
    expected = b"ttl_seconds = 300\n"

    def read_hash():
        return sha256(target.read_bytes()).hexdigest()

    class FileAdapter:
        def execute(self, request):
            target.write_bytes(expected)
            return ActionResult(
                action_id=request.action_id,
                status=ActionStatus.SUCCEEDED,
                resulting_state_hash="worker-report",
            )

    runtime = Runtime(SQLiteLedger(), state_resolver=read_hash)
    before_bytes = target.read_bytes()
    request = ActionRequest(
        action_id="edit", task_id="cache", capability="write",
        instruction="Apply the fixture TTL change.",
        precondition_hash=sha256(before_bytes).hexdigest(),
        idempotency_key="fixture-edit", adapter="fixture",
        requested_by="reviewer", actor_id="file-worker",
    )
    result = runtime.execute_action(
        request, authority=Authority(frozenset({Capability.WRITE})),
        adapter=FileAdapter(),
    )
    after_bytes = target.read_bytes()  # outside the acting adapter
    observed_hash = sha256(after_bytes).hexdigest()
    effect_matches = after_bytes == expected

The last comparison is the important line. effect_matches is a fixture-local conclusion about exact target bytes. It does not follow from the returned status, and it does not inspect every other file the worker could have changed.

Remove the adapter’s write in a copy of this example. Then make another copy that writes to a different target, and another that writes only part of the content before raising. Those variations are exactly what the pinned run in the next section executes.

Partial failure is where the record is most fragile. Until a repair that Chapter 22 will explain, CodeAI caught only RuntimeError from execution and recorded it as a failed result with a state reading; another ordinary exception could escape and leave the request without a completion. That gap mattered in the pinned run: a ConnectionError after a write let a same-key retry write again.

Current CodeAI catches ordinary adapter exceptions and records them as FAILED. It also appends action.execution_started after the precondition check and before invoking the adapter. A process death after that marker therefore no longer looks like an untouched request: it leaves no completion, but the durable record says that an effect may have happened. A resolver that itself raises after execution can still prevent the completion from being appended. Those states are reasons for reconciliation, not evidence of rollback. 5

Honest, lying, partial and wrong-target workers

The experiment for this chapter ran under a protocol frozen before execution. It starts from a real decision.recorded resting on a supported claim, applies a small patch in a disposable workspace per case, and exports the decision, request, result, invocation receipts, before and after bytes, and hashes. An independent stdlib-only verifier derives the byte comparison from the ledger and the fixture files rather than importing the producer’s success predicate.

CaseAdapter statusIntended bytes present?Runtime observation == actual bytes?
Honest adaptersucceededyesyes — report string and digest keep their origins
Lying adaptersucceededno, target untouchedyes — the recorded observation is the before-hash, exposing the discrepancy
Partial adapterfailed (ConnectionError mid-write)8-byte prefix onlyyes — same-key retry replays FAILED with 2 requests, 1 effect
Wrong-target adaptersucceededno, target untouchedyes — the fixture inventory exposes other.toml with the expected bytes
Decision basis moves latersucceeded, bytes as intendedyesyes — the decision event stays intact while standing moves basis_intact → basis_changed

Look at the lying row again, because a later experiment found something this table does not show. The runtime recorded both readings and exposed the discrepancy to anyone who compared them — and then did not compare them itself. Its own effect projection read the adapter’s succeeded and reported the effect as OBSERVED, while its two readings were byte-identical and a check of the intended post-condition returned FAIL. The word in the projection’s explanation was “observation”; what it had was a reading it never used. 7

The repair is the three-dimension model above: a reported success whose observed scope did not change is UNKNOWN and goes to a person, rather than being counted as an observed effect. The distinction that makes it honest is the one this chapter has been building — and the distinction that keeps it honest belongs to the next chapter but one:

WorkerReportsRuntime observationEffect stateVerification
Diligent — writes the intended textsucceededchangedOBSERVEDPASS
Busy — writes the wrong textsucceededchangedOBSERVEDFAIL
Idle — writes nothingsucceededunchangedUNKNOWN—
Unwatched — no resolver configuredsucceededunavailableREPORTED—

The first two rows are the load-bearing ones. Same report, same observation, same effect state, opposite verdicts. CHANGED does not mean correct; it means the scope the runtime can see is not what it was. Whether it changed the right way is a check’s question, never an observation’s. 1

The partial case exercises the exception repair described above. A ConnectionError, not a RuntimeError, is recorded as FAILED, and the retry returns the FAILED replay with the adapter invoked exactly once. The stale-basis case uses a genuine decision.recorded from an offline recorded call. Refuting evidence lands after the effect, and the standing projection changes while the old decision and the performed effect are retained.

The run’s limits stay with its result: the runtime resolver watches exactly one target file (everything else comes from the verifier-side inventory, not from any ledger field), fixtures are disposable and local with no concurrency, and process death between effect and completion is not exercised — only the caught-exception path is. The current action.execution_started recovery marker is a later implementation repair, not a result measured by this pinned run.

Whose reading is it?

Keeping the report apart from the observation has to hold field by field, and CodeAI’s own history shows how easily it does not. In the version of CodeAI that the earlier demos and the Stage 29B run used, the runtime does call its configured state resolver after an adapter returns. It does not merely ask the adapter for a hash. With no resolver, _current_state_hash() returns None. That establishes a separate origin for the reading, not independent truth: what the reading covers, how complete it is, and whether it is a good representation of the relevant world all depend on the resolver the caller configured. 8

But finalization then combined the two readings like this: 9

resulting_state_hash=result.resulting_state_hash or resulting_state_hash

The left value comes from the adapter. The right value is the reading passed into finalization by the runtime. A nonempty adapter value wins. A reader of the final resulting_state_hash cannot tell from that field alone which path supplied it; a conflicting runtime reading is not retained there. That is the defect: adapter-first finalization hides whose reading the field holds. 9

CodeAI now adds a separate observed_state_hash. Finalization sets it from the runtime reading even when the adapter supplied a different value, including a value in that new field itself. The legacy resulting_state_hash keeps its old behavior for compatibility. The older evidence bundles predate the field and do not establish this behavior; the pinned run above does exercise it. 9

Field in the revised resultHow to read it
state_hashAdapter-supplied value; its meaning depends on the adapter
resulting_state_hashLegacy adapter-first value, with runtime reading as fallback
observed_state_hashRuntime-resolver reading, or None when unavailable

Those meanings follow from the result definition and finalization; none denotes a correctness verdict. 9

Old completion records do not acquire an independent observation when reopened. Deserialization leaves the new field null if it was absent. Reusing a recorded result carries its original observation forward; it does not take a fresh reading of today’s world. 10

This separates two reports. Comparison of an expected resulting state with the observed one stays outside the contract, and an adapter’s successful status stands even when the intended file stayed unchanged. 5

An identity has a scope

The default repository resolver is a policy, not a complete description of the environment. In a repository with a resolvable HEAD, default_repository_state_hash hashes a combination of HEAD, staged diff, unstaged diff, and the sorted untracked path list. It deliberately does not read untracked file contents into that Git-based identity. 11

Consequently, changing the contents of an already-present untracked file need not change this identity. That follows from what is hashed; this chapter did not measure it. The function’s fallback path is different: it walks files and hashes paths and bytes, excluding directories and paths containing .codeai. The two paths should not be described as one universal content hash. 11

Keep three meanings separate: reported result/state is what the acting boundary reports; observed state is what a specified observer reads; expected state/effect is what the operation was intended to produce. Those concepts do not depend on an old field name having perfectly clean historical semantics.

Before comparing hashes, ask what each covers. A session identifier, a target-file digest, and a repository identity are not interchangeable. Even comparable before and after hashes answer only whether the selected representation changed. A change to the wrong file can change a repository identity; an unchanged value can be correct for an operation whose desired state already held.

For the teaching fixture, reading the target’s bytes makes the question concrete. Preserving those bytes also makes later inspection possible. A digest alone cannot show a future reader which line changed.

The real adapter has a narrower answer

OpenCodeExecutionAdapter is an existing implementation of the protocol. By source inspection, it creates a session if needed, sends the instruction through client.prompt, joins response text parts into a transcript, and returns SUCCEEDED. Its state_hash is the session ID. It records the session and transcript, leaving the resulting repository and any task-specific postcondition unread. That is read from the source; no live OpenCode run was made. 12

That is a useful transport boundary, but a session ID is not a repository hash: a prompt response and transcript can describe work without establishing the resulting file state, so the surrounding process still owes the reader an observation with a declared target and scope.

Other evidence of real effects

The available action evidence is partial — the capstone demo preserves a denied edit under READ with the file unchanged, a successful edit under WRITE attributed to human-approver, and a separate post-edit file check reported PASS. Its scope includes fake cognition, local checks, a stub state resolver, and clean reopen. It is a demo, not a pinned experiment. 13

The later capstone composition, run under a protocol frozen before execution, keeps the three origins apart on one successful edit: the adapter returned succeeded with no state of its own, the runtime recorded its resolver reading, and an independent reader matched that reading to the preserved after-bytes. The same run shows the first weakness listed below in action. Its resolver hashed one file, so when a separate branch changed a second file and then failed, the recorded observation never saw the change. 14

Stage 29B, the experiment Chapter 28 reports, has a narrower measured control. The rule-resolved answer was submitted as a WRITE effect through the action path, first under READ and then under WRITE. 15

ControlRecorded effect statusApply-adapter callsIndependently recorded target state
READdenied0File absent before and after
WRITEsucceeded1File absent before; afterwards [cache] and ttl_seconds = 300

These rows come from authority/authority-denied.json and authority/authority-granted.json, not inferred from the status labels. The latter includes the resulting text and its SHA-256. 15

This is evidence that a bounded effect happened in that control and was inspected outside its success label. The dishonest worker, partial write, wrong target, and Stage 18 decision whose basis later moves remain uncovered there. Those joined cases are covered by the pinned run earlier in this chapter.

What this is not

  • Authority belongs to Chapter 20. Requester attribution raises the question of permission; that chapter develops it.
  • Check adequacy remains open. Reading a file establishes its bytes within a time and scope; Chapter 21 asks what check can establish the required behavior.
  • Retry safety belongs to Chapter 22. The interval between an effect and a durable completion is handled there.
  • Containment and production validation are out of scope. A disposable file fixture does not establish either.

Where it is still weak

  1. Observation is configurable, and the two readings are not a transaction. A constant or incomplete resolver supplies a poor reading and the runtime does not validate its coverage; anything else may move the scope between the before and after readings, so a change is attributed to this action by proximity rather than proven to be its effect. Identical bytes do not prove that nothing happened elsewhere — which is exactly why that case is UNKNOWN. 16 1
  2. The expected outcome is outside the action contract. No built-in comparison judges the intended resulting state. 5
  3. A decision reference is optional, and it is not authority. Current ActionRequest can cite a Chapter 18 decision, and execute_action refuses a cited basis that is defeated or unknown. An action that cites no evidentiary decision can still proceed if its other gates permit it. Separately, the scheduler cannot select ACTION, so a scheduler decision cannot license a write. Evidentiary justification, scheduler governance, and authority remain distinct. 5
  4. Failure evidence can still stop at unknown. Adapter exceptions are recorded and action.execution_started marks the pre-effect boundary, but a resolver failure after execution or a process death can still leave the completion absent. The marker means the effect may have happened; it does not manufacture the missing observation. 5

project_decision_standing reports how a recorded basis differs from current claim standing, while the later project_decision_evidence projection can classify a cited basis as defeated or unknown for future use. Neither projection reverses a file write that already happened. A dedicated request.decision_id links a decision to an action when one is actually supplied; a shared task name does not invent that relationship. 17

Do this now

Thirty minutes. Take one proposed edit and follow it as far as the file.

  1. Save the original bytes in a disposable workspace. Write the desired bytes separately, before invoking a worker.
  2. Record who requested the change, who will perform it, and the expected before identity. Keep these distinct from the decision that motivated it.
  3. Run the worker, save its report, and read the target independently. Compare the bytes with the desired result and inspect other changed paths.
  4. Replace the worker with one that only says “done.” Identify which record now exposes the discrepancy and which record still says success.

If you are building with an assistant:

Make one disposable file action inspectable from request to effect.
Record the request before execution. Keep the worker's status and state
report separate from a runtime-controlled reading of the target.
Preserve before and after bytes and declare the hash policy.
Do not infer correctness from status or from a changed hash.
Exercise honest, no-op, partial-write, and wrong-target workers.
Count invocations outside the completion record and inventory changed paths.
Show missing observations as missing. Preserve failures without inventing
a successful outcome. Leave authority, check adequacy, and retry policy
as explicit subsequent work.

Failure modes

  • Treating a decision as an effect. A well-supported choice can still leave the file untouched.
  • Treating “done” as an observation. A worker’s report is evidence of what it reported.
  • Comparing hashes without their policies. Different representations can answer different questions.
  • Calling every changed hash correct. The wrong target can change too.
  • Reading failure as rollback. A worker can fail after making part of its change.
  • Rewriting history when the basis moves. New knowledge changes what is justified now, not what happened then.

What this chapter established

  • Decided ≠ requested ≠ performed ≠ observed. Each answers a different question, and even an observed state still has to be judged against what was intended.
  • A worker’s “done” is a report, not an observation. Keep the worker’s status and state report apart from the runtime’s resolver reading. A separate field preserves origin; a declared scope says what that reading can actually establish.
  • Know what a hash covers. A session ID, a file digest and a repository identity answer different questions. CHANGED means the selected representation changed, not that the intended change occurred. UNCHANGED does not rule out an effect outside the resolver’s scope.
  • Failure is not rollback. A worker can fail after changing part of its target. Current CodeAI records action.execution_started before the adapter so a later missing completion can project to an unknown effect rather than “nothing happened”; that recovery seam does not make the missing observation appear.
  • A decision’s basis can move after the effect. New evidence changes what is justified now, not what was done. A current action may cite a decision whose basis is checked before execution, but neither that check nor a later standing projection undoes an effect already performed.

What CodeAI showed. The pinned five-case run exercised honest, lying, partial and wrong-target workers and a decision whose basis moved later. The adapter reported succeeded for the honest, lying and wrong-target workers alike, while the configured runtime observation matched the watched target’s actual bytes in every case, and a separate verifier recomputed each row from the ledger and fixture files. Stage 29B’s authority control adds a measured file write with a separately recorded result. 15

The implementation has since added seams that the pinned run did not measure: observed_state_hash keeps the runtime reading separate from the adapter’s state report; action.execution_started marks when an effect may have begun; and an optional action decision_id can be checked against current decision evidence before execution. The resolver in the pinned run still watched one file, process death was not exercised, and comparison with the intended resulting state remains outside the action contract.

Evidence notes

The pinned run. The decision-to-effect bundle is experiments/applied-ai/evidence/decision-to-effect/2026-09-14-a1b562a/. Four seeded corruptions separate a broken record from a poorly described effect — an altered observed hash, a worker report substituted for the observation, a removed changed-path inventory entry, a rewritten decision basis — each rejected with the failure named. The exception repair that the partial case exercises was uncommitted when the run executed; its hashed files match the repair as later committed.

Stage 29B. The Stage 29B verifier checks the authority control within its wider experiment. The wider bundle is not wholly green: its report preserves a failed spend claim, L-D, caused by usage lost on an empty-output failure path. The action control should be cited locally, not used to describe the whole bundle as passing. 18

Next

A callable worker gives the process the ability to change something. An observation gives it a way to learn what was left behind. Neither answers who was allowed to request the change.

Continue with Capability Is Not Authority.

References

  • Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing Reasoning and Acting in Language Models. ICLR, 2023. arXiv:2210.03629.
  • Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. PAL: Program-aided Language Models. ICML, PMLR 202:10764–10799, 2023. Proceedings.

Implementation sources: the pinned Chapter 19 evidence preserves the historical action path and its observation/exception repairs; current-source claims above were checked against their successors rather than projected backward into that run. src/codeai/adapters.py: ActionRequest, ActionResult, ActionStatus, ExecutionAdapter; src/codeai/runtime.py: Runtime.execute_action, _finalize_action_result, _action_result_from_payload, _append_action_result_event, _current_state_hash, default_repository_state_hash; src/codeai/actions.py: EffectState, StateObservation, ActionWorkState, project_action_state; src/codeai/opencode.py: OpenCodeExecutionAdapter.execute; src/codeai/evidence.py: project_decision_standing, project_decision_evidence; src/codeai/governance.py: the later operation-governance boundary that keeps ACTION outside scheduler-selectable operations. Regression source: tests/test_action_observation.py; related current recovery and evidence tests are separate later work. Evidence: experiments/applied-ai/evidence/capstone/; capstone-composition/; execution-ladder/2026-09-14-7a0d43b/, especially authority/; working-state/2026-09-13-68f4ba0/; decision-to-effect/2026-09-14-a1b562a/ (pinned five-case run with a separate verifier). Footnote prefixes: s inspected source, m pinned run, d preserved demo, r frozen report.


  1. Measured run: experiments/W1-R1-effect-observation.md and its recheck; src/codeai/actions.py (EffectState, StateObservation), example examples/applied_ai/ch19_effect_observation.py, seam limits in docs/seams/action-recovery.md. ↩︎ ↩︎ ↩︎

  2. Source inspection: src/codeai/adapters.py (ActionRequest). ↩︎ ↩︎

  3. Interface inspection: src/codeai/adapters.py (ExecutionAdapter). ↩︎

  4. Source inspection: src/codeai/adapters.py (ActionResult, ActionStatus). ↩︎

  5. Source inspection: src/codeai/runtime.py (Runtime.execute_action). ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  6. Source inspection: src/codeai/runtime.py (Runtime.execute_action, Runtime._finalize_action_result). ↩︎

  7. Measured run: CodeAI’s frozen Wave 1 composition audit, probe D (experiments/W1-composition-results.md, baseline runtime f7d2910). Chapter 29 reports the audit in full. ↩︎

  8. Runtime inspection: src/codeai/runtime.py (Runtime._current_state_hash, Runtime.execute_action). ↩︎

  9. Finalization inspection: src/codeai/runtime.py (Runtime._finalize_action_result). ↩︎ ↩︎ ↩︎ ↩︎

  10. Source inspection: src/codeai/runtime.py (Runtime._action_result_from_payload, Runtime.execute_action). ↩︎

  11. Source inspection: src/codeai/runtime.py (default_repository_state_hash). ↩︎ ↩︎

  12. Adapter inspection: src/codeai/opencode.py (OpenCodeExecutionAdapter.execute). ↩︎

  13. Unpinned demonstration: experiments/applied-ai/evidence/capstone. ↩︎

  14. Measured run: experiments/applied-ai/evidence/capstone-composition. ↩︎

  15. Measured run: experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b. ↩︎ ↩︎ ↩︎

  16. Source inspection: src/codeai/runtime.py (Runtime._current_state_hash). ↩︎

  17. Source inspection: src/codeai/evidence.py (project_decision_standing). ↩︎

  18. Report: experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/chapter-evidence-report.md. ↩︎