← Applied AI

Retries Are Side Effects Too

Retry is a new effect, replay is a returned record, and a duplicate is the failure to tell them apart. An idempotency key may replay only the same recorded operation, never bypass current authority — and a failure status never proves the effect did not happen.

Part 4 — Make It Safe and Verifiable

Retry, replay, duplicate

Retrying is the most ordinary recovery move in software, and in an AI application it is often the most dangerous one. When a call or an action fails, or only seems to, the loop tries again. If the first attempt already sent the email, charged the card, wrote the file or ran the model, the retry does it twice. A timeout says nothing about whether the far side acted.

retry  ≠  replay  ≠  duplicate

A retry is a new attempt to cause an effect. A replay returns a result already recorded, instead of acting. A duplicate is the same effect caused twice: the failure the other two exist to prevent. By the end of this chapter you will be able to say what evidence justifies returning a recorded result instead of acting again, why an idempotency key alone is not that evidence, why current authority must be checked even on a replay, and why a failure status never proves that the effect did not happen.

The line that may already be there

Picture an adapter that appends one line to a file. The process dies before the completion is recorded. On restart the ledger shows no completed result — but the file has changed.

Three responses are available, and two of them lie. Appending again may write the line twice. Writing a success record by hand manufactures evidence for an effect nobody observed. Saying “unknown” is the only honest answer, and it is also the least actionable one. The question that matters is narrower than “what do we do?” It is: what evidence would justify another append?

Chapter 21 ended with failure inviting another attempt. This chapter controls the invitation. A retry is not a loop. It is a claim about the first attempt — what it intended and observed, whether it was authorized, what was checked — and a decision about what may happen to the world next.

When may the runtime repeat work without duplicating an effect?

Retry, replay, duplicate

Each term names a different thing that can happen when a request arrives twice:

TermMeaningEffect on the world
RetryA new attempt to cause an effectMay cause another real effect
ReplayA previously recorded outcome returned instead of actingNo new effect by construction
DuplicateThe same intended effect physically caused more than onceThe failure the other two must prevent

An idempotency key alone does not establish which of these is safe. A key is a lookup handle, not the identity of the intended operation. The chapter’s operational question is therefore: what evidence justifies returning a prior result instead of taking another effect?

The distributed-systems literature gives this question its shape. RIFL converts at-least-once RPCs into exactly-once ones by durably recording completed results and returning the saved result when a retry arrives, and it names four problems every such mechanism must solve: RPC identification, completion-record durability, retry rendezvous, and garbage collection (Lee et al., 2015). Its hardest requirement is that the completion record be created atomically with the operation’s mutations: in the paper’s words, “without a visible completion record, or vice versa.”

That atomicity is exactly what CodeAI’s file-plus-ledger path cannot provide — the adapter writes wherever it likes, and the ledger hears about it afterwards. The mapping of RIFL’s four problems onto CodeAI’s action path is the book’s own, and the gap is the point: a key lookup is one quarter of RIFL, not RIFL.

Birrell and Nelson’s RPC design is the older foundation: remote calls need retransmission because responses get lost, and a timeout says nothing about whether the far side acted (Birrell and Nelson, 1984). That supports separating the logical request from its attempts — CodeAI’s Chapters 11 and 16 already do — but retransmission alone is at-least-once by construction. It does not, by itself, tell a retry from a duplicate. The application to CodeAI’s replay rule is again the book’s mapping, not the paper’s result.

ARIES is the deliberate contrast, used lightly. Write-ahead logging recovers database transactions because the log is the recovery path for pages the system manages (Mohan et al., 1992). Arbitrary file writes and external API effects live outside those managed pages. No logging discipline around them manufactures atomicity. The chapter borrows ARIES’ vocabulary — completion records, recovery obligations — and none of its guarantees.

The rule, and one implementation of it

The retry rule is short: a current request must survive its pre-effect gates before prior history is disclosed, and a same-key completion may be replayed only when the incoming operation has the same recorded identity. Current execute_action has accumulated later book seams around that rule. After action.requested it records operation governance, resolves current authority, and, when the request cites an evidentiary decision, checks that basis. Only then does it look for a prior completion by idempotency key. On a hit, _replay_action_result checks the action fingerprint before returning history. On a miss, the fresh path checks the precondition, appends action.execution_started, invokes the adapter, reads the state resolver again, and appends action.completed. 1

The replay branch is therefore not a bare key lookup. The current request has to pass the gates that apply now; only then may identity match and replay return an old result. 2

Authority first. If the action names a directive, current CodeAI resolves authority from that recorded directive chain; if it names none, it falls back to the caller-supplied Authority and records that basis. A refusal is recorded as action.authorization_refused plus a DENIED completion before replay lookup, so the denied requester does not learn whether the key already exists or which fingerprint dimensions it would match. Ability to replay an old record is not authority to request the operation now. 1

Identity second. The incoming request’s fingerprint is compared against the fingerprint rebuilt from the original request’s recorded action.requested payload, so the check reproduces after ledger reopen. On mismatch the runtime appends action.replay_refused — naming the key, the original action, the fingerprint version, and the mismatching dimensions — and raises IdempotencyConflictError before any effect. A conflict leaves durable evidence, not only an in-memory exception; the bare raise would have left action.requested unresolved, the same class of gap Chapter 21 closed for checks. 3

Replay last. Only the same key plus the same identity returns the recorded completion, relabeled with the new action ID and reused_from_action_id pointing at the original physical effect. Whatever the first execution recorded — success, failure, even denial — is what replay returns. 2

The fingerprint, action-fingerprint-v1, is deliberately narrower than the request and wider than the key: 4

In the fingerprintWhy
capabilityA different permitted operation is a different operation
instructionChanged directions change the intended effect
payload (canonical SHA-256)Application operation data; key order ignored, values hashed
precondition_hashThe world state the operation was planned against
Effective adapter IDWho performs it is part of what was requested
Out of the fingerprintWhy
action_id, task_idThis request instance and its correlation, not the operation; mirrors the call fingerprint excluding task identity
directive_id, requested_by, actor_id, decision_idAuthority, attribution and any cited evidentiary decision are evaluated on the current request rather than frozen into the replay identity

The call path is the architectural precedent, and the comparison is instructive. Call replay looks up call.completed events that carry a matching call.manifest — legacy completions without terminal evidence are misses — then _check_replay_fingerprint compares prompt hash, chamber, requested model, and rendered context, raising IdempotencyConflictError naming the dimensions before any provider effect. An identical re-request replays with replayed: true and the original call ID. 5

Two differences matter. First, the call conflict stays in memory: call.requested is appended, the check raises, and no durable conflict event follows. The action path now records its refusals; the call path’s gap is noted, not repaired here. Second, calls carry a manifest — a runtime-built record of what was actually sent — while actions have only the request the caller supplied. The action fingerprint is therefore weaker by construction: it binds the requested operation, not an independently built manifest of the performed one. 6

The precondition rule, stated plainly

Precondition is part of request identity, and it is not re-evaluated before returning a recorded completion. Both halves need their reason on the page.

Identity, because the precondition names the world the operation was planned against. A request planned against state A and a request planned against state C are different planning instances even when their instructions match; conflating them lets one worldview borrow another’s outcome.

No re-evaluation, because re-checking would destroy legitimate idempotent replay. The original precondition was state A; the action moved the world to B; the same request arrives again. If replay required the current state to still equal A, then a successful action would destroy the condition required to replay its own result — the mechanism would fail exactly when it worked.

The current-state comparison belongs to fresh execution, where it refuses drift before acting. Replay answers a different question — “what did this operation instance record?” — and the recorded precondition is part of how the instance is identified, not a gate on reading its history. 7

The cost of that choice is explicit: if the world moved back to A after a drift failure, the same key still replays the old FAILED rather than re-evaluating. A changed context that should produce a new decision needs a new key. Keys scope one logical instance and its recorded outcome; they are not subscriptions to the world.

The demo, read exactly

The preserved retry-effects demo runs five offline cases with fake adapters, and its independent verifier checks the summary plus a seeded corruption. 8

CaseWhat ranRecorded result
ASame key and payload twice2 requests, 1 physical effect, second carries reused_from_action_id: a1
BSame key, different instruction and capability, empty authoritySilently replayed the old success, 0 new effects — the historical defect
CSame call key, different promptIdempotencyConflictError naming prompt_hash, 0 new effects; identical re-request replays with replayed: true
DAdapter applies the effect, then raises RuntimeErrorStatus failed with the error preserved though the file changed; retry replays failed, 0 new effects
ERequested state differs from current statefailed with precondition mismatch, 0 effects

Case B is the defect the rule above prevents, and the evidence discipline matters here: the demo ran against the code from before the repair, so it establishes that the defect existed, not that the repair works. The repair is source-inspected and regression-tested in current CodeAI — and the pinned rerun below replays every demo case against the repaired semantics with a ledger-based verifier. Historical demo found the defect; current source contains the repair; regression tests exercise it; the rerun measures it. 8

Under the current rule, case B’s two variants separate cleanly. Same key with a different operation and proper authority → IdempotencyConflictError naming the dimension, action.replay_refused on the ledger, zero effects. Same key with the same operation but no authority → DENIED completion, zero effects, no information about the recorded operation. Both were executed as teaching code, not as a pinned run. 2

Failure is cached, and that is a decision

Case D deserves its own section because it breaks the most natural retry loop. The adapter wrote the file and then raised. The ledger says FAILED. A naive retry — catch the exception, call again with a fresh key — writes the file twice. A same-key retry replays the FAILED without touching the file. Both behaviors are now covered, and neither recovers anything: 8

write_file()
raise RuntimeError("connection lost")

A fragment with the real APIs, executed as teaching code. The file marker appears once; the ledger says failed; the retry returns the failure with reused_from_action_id set and the adapter untouched. 1

first = runtime.execute_action(req("d1", "kc"), authority=auth, adapter=crasher)
second = runtime.execute_action(req("d2", "kc"), authority=auth, adapter=crasher)
assert (first.status, second.status) == (ActionStatus.FAILED, ActionStatus.FAILED)
assert second.reused_from_action_id == "d1"
assert crasher.calls == 1

This gives the chapter’s hardest-won distinction: operation reported failed ≠ effect did not happen. Replaying a recorded failure says “this request already produced this recorded outcome.” It returns history without attempting recovery. In the historical demo, the post-failure state reading was the available runtime witness. Current CodeAI adds a stronger recovery record: action.execution_started is committed before the adapter runs, with the runtime’s pre-effect reading when available. A failure after that marker projects the effect as UNKNOWN even though the reported result is FAILED. Later reconciliation may add evidence about the effect without rewriting the failure. 1

Where the cache stops

Caching failure protects only failures the ledger hears about, and that boundary was narrower than case D made it look.

Case D’s adapter raised RuntimeError, and before the repair described below, that was the only adapter exception execute_action turned into a recorded FAILED. Real transport failures are usually something else: ConnectionError and TimeoutError are OSError subclasses. A probe during editing gave an adapter that writes a marker and then raises ConnectionError. The exception escaped, the ledger held action.requested with no completion, and a same-key retry found no record to replay — so it called the adapter again. Two adapter calls, two markers, the exact duplicate this chapter exists to prevent. 1

The repair, since committed to CodeAI, is one clause: any adapter Exception after the authority and precondition checks is recorded as FAILED with the runtime’s observation. A regression test repeats the probe’s shape and asserts one adapter call, two FAILED completions, and a replay link. Rerunning the probe after the change gives one call and one marker. Neither the probe nor the test is a pinned stage run, and the historical demo is unchanged. 9

The pinned Stage 22 rerun froze this opening scene in its then-current form: if the process died after the effect but before action.completed, there was no completion to replay and no action-level recovery mechanism. That remains the measured result for that stage. Current CodeAI has since inserted a durable action.execution_started marker before the adapter is invoked, so the same crash gap is now visible as an unresolved effect rather than indistinguishable from a request that never reached execution. 10

The diagram below is the pinned rerun’s historical timeline — recorded completion replays, while the unrecorded crash gap could repeat:

    sequenceDiagram
    participant Caller
    participant Runtime
    participant Adapter
    participant World
    participant Ledger
    Caller->>Runtime: execute_action (key K)
    Runtime->>Ledger: append action.requested
    Runtime->>Adapter: execute
    Adapter->>World: effect happens
    alt completion recorded
        Adapter->>Runtime: return result
        Runtime->>Ledger: append action.completed
        Caller->>Runtime: retry same key K
        Runtime->>Ledger: lookup K → replay, no new effect
    else crash before the record
        Adapter--xRuntime: process dies
        Caller->>Runtime: retry same key K
        Runtime->>Ledger: lookup K → nothing
        Runtime->>Adapter: execute again
        Adapter->>World: effect happens twice
    end
  

Current CodeAI now has an action-level recovery projection. project_action_state reads a current-version action.requested with no execution marker as effect NONE and START_ACTION; action.execution_started with no completion becomes effect UNKNOWN and RECONCILE_EFFECT. Records written before the execution marker existed stay UNKNOWN rather than being read as safe. open_effects lists those unresolved actions. 10

reconcile_action appends later evidence without editing the original result. It records effect_confirmed, no_effect_confirmed, or still_unknown; a confirmed no-effect can make START_ACTION the next operation, while still_unknown remains a legitimate final answer. The actor, basis and evidence references are preserved, but the runtime does not independently validate those evidence references. 10

This repairs observability, not atomicity or exactly-once execution. _find_action_result still replays only completed actions. A caller that ignores the UNKNOWN projection and directly submits the same key again can still find no completion and execute again. The recovery projection tells the process not to guess; it is not a lock around execute_action. These recovery APIs and tests are later current-source work, not results retroactively measured by the pinned Stage 22 rerun. 11

Two threads, one key, two effects

Sequential duplicate safety does not establish concurrent safety, and the ledger shows why: appends commit individually, the events table constrains only event_id uniqueness, and no lock, key uniqueness, or compare-and-set spans the lookup-to-append interval. Two requests can both find no completion, both execute, and both record. 12

A bounded threaded probe confirms the window is real, not theoretical: two threads submitting the same key against one ledger file, three runs, two physical effects and two successes every time. An earlier variant of the probe, with threads opening the ledger simultaneously, failed even sooner with database is locked — concurrent writers are unsupported at the connection level here, not merely racy.

That probe is now a frozen diagnostic inside the pinned rerun: barrier-synced, race window held open, 3 runs with 6 submissions and 6 effects preserved, repaired nothing. The durable claim remains the source-inspected absence of coordination, and the chapter claims nothing about concurrent safety. Any future concurrency mechanism will be measured against those six effects. 12

What this is not

  • Not exactly-once. Sequential replay suppresses known duplicates; concurrent submissions demonstrably duplicate, and crash-gap effects stay unknown rather than counted.
  • Not transactional. No atomicity spans adapter effect and ledger append — RIFL’s central requirement is absent, not approximated.
  • Retry policy lives elsewhere. The runtime returns records and refuses conflicts; deciding when trying again is wise belongs to policy over the evidence Chapters 18–21 preserve.
  • Completion records accumulate. RIFL’s fourth problem — leases, acknowledgments, reclamation — has no counterpart here.
  • A request binding, not a manifest. The fingerprint binds the requested operation, not an independently built record of what was performed.

Where it is still weak

  1. Concurrent duplicates are unprotected. No lock, no key constraint, no compare-and-set spans the check-to-act interval; the frozen probe shows two effects from two threads. 12
  2. Completion and effect are not atomic. action.execution_started makes the crash gap visible, but it does not make the effect and action.completed one transaction. 13
  3. Recovery is explicit but not self-enforcing. project_action_state, open_effects and reconcile_action can say UNKNOWN and record an operator reconciliation, but execute_action does not consult an unresolved same-key action before fresh execution. The reconciliation’s evidence references are recorded, not independently validated by the runtime. 14
  4. Call conflicts leave no durable conflict record. Action refusals append action.replay_refused, whereas a call fingerprint conflict still raises with the request history but no equivalent call.replay_refused. 15
  5. Completion records are never reclaimed. Every replayable outcome lives forever; no acknowledgment, lease, or retention rule exists. 11
  6. A missing current state still skips the precondition comparison on the fresh path. “Precondition supplied” remains weaker than “precondition enforced.” 1
  7. DENIED outcomes can still replay. The new request must pass current authority first, but once it does, the same key can return the earlier DENIED completion. A changed authorization context that deserves a new attempt needs a new key. 2
  8. The evidence stages must stay separate. The pinned rerun establishes the repaired sequential replay semantics and the six-effect concurrency baseline for its version. action.execution_started, action-state projection and reconciliation are later current-source repairs, exercised by source-level tests rather than retroactively added to that run.

Do this now

Thirty minutes. Retry something that already ran, and prove you did not do it twice.

  1. Run an effectful action with a fixed key and a counting adapter. Submit the identical request again and confirm one effect, two completions, and the reuse link — then reopen the ledger and confirm it again.
  2. Resubmit the key with a changed instruction and watch the conflict name the dimension. Confirm zero new effects and find the refusal event. Then resubmit with the same operation but no authority and confirm a denial that reveals nothing.
  3. Build a crash-after-effect adapter. Confirm FAILED with the effect present, retry with the same key, and write down exactly what the ledger does and does not establish about the world.
  4. Race two threads on one key in a scratch directory. Count the effects. Write the number down next to the sequential result from step 1.

If you are building with an assistant:

Make repetition explicit: retry is a new attempt, replay returns a recorded
outcome, duplicate is the same effect twice. Bind each idempotency key to one
intended operation with a versioned fingerprint of capability, instruction,
payload, precondition, and performer; refuse collisions durably before any
effect. Check current authority before returning any recorded result, and let
denials reveal nothing about what is recorded. Do not re-evaluate the original
precondition on replay; do re-check it before fresh execution. Cache failures
as history, not as proof of no effect, and leave crash-gap reconciliation to
an explicit operator decision. Never claim exactly-once, transactions, or
concurrent safety without evidence.

Failure modes

  • Minting fresh keys for old operations. A loop that creates new keys on every attempt manufactures duplicates, not safety.
  • Treating the key as the operation. A lookup handle does not establish that two requests meant the same thing.
  • Letting replay skip authority. A prior success is not permission for the current caller.
  • Reading FAILED as “nothing happened.” A failure record is history, and the effect may already be real.
  • Mistaking replayed failure for recovery. Returning the old FAILED did not fix anything; it only refused to make it worse.
  • Re-checking preconditions on replay. Requiring the old world to still hold destroys the replay of the action that changed it.
  • Assuming sequential safety covers concurrency. One key, two threads, two effects — the probe number goes here, not the hope.

What this chapter established

  • Retry ≠ replay ≠ duplicate. A retry may cause a new effect, a replay returns an outcome already recorded, and a duplicate is the failure both exist to prevent. A key alone establishes none of them.
  • Bind the key to the operation. An idempotency key is a lookup handle. What justifies returning a recorded result is a versioned fingerprint of the intended operation — capability, instruction, payload, precondition, performer — and a collision has to be refused durably before any new effect.
  • Check current authority before returning history. A prior success is not permission for the requester in front of you now. Current CodeAI resolves a named directive from the record, or records caller-supplied authority as a fallback when no directive is named, before it discloses a replay.
  • Identify with the precondition; do not re-evaluate it on replay. The state an operation was planned against is part of which instance it was. Re-checking it on replay would destroy replay of the action that changed the world. A context that deserves a new decision deserves a new key.
  • Failed ≠ ineffectual. A failure record is history, not proof that nothing happened. Current CodeAI also marks execution before the adapter runs, so a missing completion after that marker projects to an UNKNOWN effect rather than to “safe to retry.” Reconciliation can add later evidence without rewriting the original result.

What CodeAI showed. The historical demo preserves the key-only replay defect, fingerprint-bound call conflicts, cached post-effect failures and drift refusal as they were observed. 8 The pinned Stage 22 rerun measured the repaired sequential semantics: exact duplicates replayed with one effect, key collisions produced durable refusals with zero second effects, current authority was checked before replay, failures replayed without another effect, reopen preserved replay, and the deliberately widened concurrency window produced 6 submissions and 6 effects. It did not establish crash-gap recovery or concurrent safety.

Current CodeAI goes further in source and tests: action.execution_started marks the pre-effect boundary; project_action_state and open_effects expose UNKNOWN effects; reconcile_action records effect_confirmed, no_effect_confirmed, or still_unknown. Those later mechanisms improve recovery evidence but do not make the effect atomic with the ledger, do not coordinate concurrent same-key callers, and do not stop a direct caller from bypassing the UNKNOWN projection and invoking the action again. 10 12

Evidence notes

Independent verification. The demo’s independent verifier imports no CodeAI code. It requires A at 2 requests / 1 effect with the reuse link, B replaying silently (asserted as the limitation, not as success), C raising with prompt_hash named and zero effects, D caching the failure with zero new effects, and E rejecting drift — plus a mutated copy claiming two physical effects, which it must reject. 16

Its limit mirrors the earlier chapters’ verifiers: it checks summary fields, not the ledger, so a summary that miscounted effects consistently would pass. The pinned rerun in experiments/applied-ai/evidence/retry-rerun/2026-09-14-1b3c7a2/ answers that weakness with a frozen protocol, no fixes during the run, and a stdlib-only verifier recomputing from durable evidence — request counts, invocation receipts, marker files, fingerprints, keys, authority decisions, replay provenance, original and replayed IDs, reopen identity — plus five seeded corruptions (altered replay provenance, an inflated receipt, a dropped refusal or request, a rewritten FAILED completion), each rejected with the failure named.

The rerun covers exact sequential duplicates (2 requests, 1 effect, reuse link resolves); key collisions with zero second effects and the preserved conflict naming instruction; authority changes that return a DENIED completion with no reuse and no new effect; duplicates across process reopen; crash-gap cases for both exception families showing what the record establishes (FAILED replay, one effect) and what it cannot (whether the bytes changed is witnessed only by the marker file, and recovery stays operator judgment); precondition drift refused before effect.

It also freezes a diagnostic concurrency probe — barrier-synced threads with the race window held open, recorded without repair: 3 runs, 6 submissions, 6 effects, every row reused=None. That probe is the frozen baseline the concurrency work must improve on.

Next

Repetition is now explicit: completed identical requests can return records, materially different same-key requests conflict, unauthorized requests are denied before replay disclosure, and unresolved effects stay explicit until reconciliation either resolves them or records that they are still unknown. With acting, checking, and repeating all recorded, the process can entertain more than one proposal at a time — provided they cannot see each other.

Continue with Blind Before You Compare.

References

  • Collin Lee, Seo Jin Park, Ankita Kejriwal, Satoshi Matsushita, and John Ousterhout. Implementing Linearizability at Large Scale and Low Latency. SOSP, 2015. DOI.
  • Andrew D. Birrell and Bruce Jay Nelson. Implementing Remote Procedure Calls. ACM Transactions on Computer Systems 2(1):39–59, 1984. Publication page.
  • C. Mohan, Don Haderle, Bruce Lindsay, Hamid Pirahesh, and Peter Schwarz. ARIES: A Transaction Recovery Method Supporting Fine-Granularity Locking and Partial Rollbacks Using Write-Ahead Logging. ACM Transactions on Database Systems, 1992. Publication page.

Implementation and evidence sources: the historical retry-effects demo and the pinned Stage 22 rerun remain evidence about the versions that produced them; later current-source repairs are distinguished rather than projected backward into those bundles. Current replay and execution ordering were checked in src/codeai/runtime.py: execute_action, _replay_action_result, _check_action_replay_fingerprint, _find_action_result, _action_fingerprint, ACTION_FINGERPRINT_V1, IdempotencyConflictError, invoke_recorded_call, _check_replay_fingerprint, _find_recorded_completion; current action recovery in src/codeai/actions.py: ACTION_RECOVERY_V1, project_action_state, project_open_effects, reconcile_action; authorization in src/codeai/authority.py and Runtime.authorize_action; later operation-governance and decision-evidence gates in src/codeai/governance.py and src/codeai/evidence.py; request/result types in src/codeai/adapters.py; ledger behavior in src/codeai/ledger.py. Relevant current tests include tests/test_action_replay.py and tests/test_action_recovery.py. Teaching fragments were illustrative code, not pinned stage runs; the threaded probe is frozen inside the pinned rerun. Evidence: experiments/applied-ai/evidence/retry-effects/, retry-rerun/2026-09-14-1b3c7a2/ (frozen protocol, producer, separate stdlib verifier, five seeded corruptions rejected), and working-state/2026-09-13-68f4ba0/. Footnote prefixes: s inspected source, m pinned run, d preserved demo, r frozen report. The retry-effects bundle and pinned rerun are unchanged by the later action-recovery source revisions.


  1. Source inspection: src/codeai/runtime.py (Runtime.execute_action). ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Source inspection: src/codeai/runtime.py (Runtime._replay_action_result). ↩︎ ↩︎ ↩︎ ↩︎

  3. Source inspection: src/codeai/runtime.py (Runtime._check_action_replay_fingerprint). ↩︎

  4. Source inspection: src/codeai/runtime.py (_action_fingerprint). ↩︎

  5. Source inspection: src/codeai/runtime.py (Runtime.invoke_recorded_call, Runtime._check_replay_fingerprint). ↩︎

  6. Source inspection: src/codeai/runtime.py (Runtime._find_recorded_completion). ↩︎

  7. Source inspection: src/codeai/runtime.py (Runtime.execute_action, Runtime._replay_action_result). ↩︎

  8. Unpinned demonstration: experiments/applied-ai/evidence/retry-effects. ↩︎ ↩︎ ↩︎ ↩︎

  9. Source inspection: tests/test_action_replay.py (test_non_runtime_error_after_effect_is_recorded_and_replayed_not_repeated). ↩︎

  10. Source inspection: src/codeai/actions.py (ACTION_RECOVERY_V1, project_action_state, project_open_effects, reconcile_action) and src/codeai/runtime.py (Runtime.execute_action). ↩︎ ↩︎ ↩︎ ↩︎

  11. Source inspection: src/codeai/runtime.py (Runtime._find_action_result). ↩︎ ↩︎

  12. Source inspection: src/codeai/ledger.py (SQLiteLedger.append); measured concurrency baseline: experiments/applied-ai/evidence/retry-rerun/2026-09-14-1b3c7a2/. ↩︎ ↩︎ ↩︎ ↩︎

  13. Source inspection: src/codeai/runtime.py (Runtime.execute_action, Runtime._find_action_result). ↩︎

  14. Source inspection: src/codeai/actions.py (project_action_state, project_open_effects, reconcile_action) and tests/test_action_recovery.py. ↩︎

  15. Source inspection: src/codeai/runtime.py (Runtime._check_replay_fingerprint). ↩︎

  16. Unpinned demonstration: experiments/applied-ai/evidence/retry-effects/verify_retry.py. ↩︎