← Applied AI

A Successful Call Is Not Finished Work

When is AI work actually done? A model call can succeed without the task succeeding, and a check can pass without the task being complete. Generation, checking, acceptance and completion are four different facts owned by different parts of the process. Built in CodeAI as explicit acceptance, interrupted by a killed process, reopened, and attacked twelve ways.

Part 2 — Get the Model Out of the Chat Box

The model is not the process

An AI pipeline gets a response back, the response looks right, and something marks the work done. That step is where many AI-enabled processes quietly lose track of the truth.

A successful model call tells you the call worked. It does not tell you that anyone checked the output against what the task required, that the thing checked is the thing being used, that someone with standing agreed, or that the agreement was written down. Collapse those into one status field and “done” starts to mean “the model returned something plausible”.

This chapter separates four facts that usually share that field:

generation succeeded  ≠  artifact checked  ≠  artifact accepted  ≠  task completed

The output belongs to the call. “Done” belongs to the task.

What facts must exist before work can be called complete?

By the end of the chapter you will be able to answer that for any piece of AI work: the exact artifact, a check that ran on those bytes against the task’s own criteria, an acceptance by a role other than the producer, and a completion that points back to that acceptance. You will also see why the answer holds when a process is killed, when someone clicks twice, and when someone edits the output after it was checked.

Four facts that are not the same fact

Each of those four facts is owned by something different:

Task
 │
 ├── Call ─────────── generated artifact          (what the model produced)
 │
 ├── Check ────────── evidence about that artifact (what a deterministic procedure observed)
 │
 ├── Acceptance ───── cites the call, the artifact and the check,
 │                    made by an eligible role     (a decision)
 │
 └── Completion ───── valid only when caused by that acceptance

Between the output and “done” lies the process. The model produces cognition inside a task that the runtime still owns:

    flowchart TD
    subgraph TASK["task — owned by the runtime, not the model"]
        direction TB
        C(["model call<br/><i>produces cognition</i>"]) --> AR["artifact<br/><i>what the model produced</i>"]
        CK["check<br/><i>deterministic observation</i>"] --> EV["evidence<br/><i>about that artifact</i>"]
        AR --> AC["acceptance<br/><i>cites call + artifact + check</i>"]
        EV --> AC
        AC --> CO["completion<br/><i>valid only when caused by acceptance</i>"]
    end
  

The model is not the process, and that is not a complaint about models. A model call can produce the first artifact in that diagram, but the producing call cannot establish the other three facts. A later model could be shown a check record or an acceptance, but its statement about them would still be another claim unless the process record supported it. Whether the exact bytes were checked, which bytes the checker saw, whether an eligible role accepted them, and whether that acceptance was recorded are facts about the process around the call. The runtime has to hold them.

A candidate that looks right, an unfinished task

Here is what that looks like on a small task: repair one paragraph.

The new cache makes every page load 73% faster, according to the platform team [S1].
It stores rendered fragments close to readers.

It had three written criteria: the paragraph must contain no percentage figure, the source marker [S1] must appear exactly once, and the second sentence must survive word for word.

The call went through the recorded path from Chapter 11 and came back with this:

The new cache is intended to make page loads faster, according to the platform team [S1].
It stores rendered fragments close to readers.

The candidate appears to satisfy all three written repair criteria. The dialect was handled, generation finished normally, and the attempt was interpreted and preserved. None of that is task acceptance. A new process reopened the ledger and asked about the task:

call_status:        succeeded
generation_state:   complete
task_completion:    incomplete   (no acceptance recorded)

Nothing went wrong. The task is still not complete.

That is the line Chapter 11 ended on, task_status: not automatically completed, and every chapter since has carried it forward. Chapter 1 called HTTP 200 a transport fact, not a work outcome. This chapter makes that distinction executable.

The check belongs at the end

In 1984, Saltzer, Reed and Clark worked through a problem they called careful file transfer (Saltzer, Reed and Clark, 1984). A file has to move from one computer’s disk to another’s without damage. Many things can fail along the way. The disk can return bad data. The software can make a buffering mistake. Memory can flip a bit. The network can drop, alter or duplicate packets. A host can crash partway.

You could make every layer more reliable. Their point was that this cannot finish the job. The transfer is trustworthy only when the application at the far end reads back what was actually stored, compares a checksum with the original, and commits only if they agree. A reliable network removes one threat and leaves the rest. In their words, the end-to-end check of the file transfer application must still be implemented, however reliable the lower layers become.

They also told a “too-real” story. A network at MIT checksummed every hop between gateways. One gateway had a fault that occasionally swapped a pair of bytes while copying between buffers, where no checksum covered the data. Source files passing through it were corrupted, and their owners were left with the ultimate end-to-end check: comparing against old listings by hand.

This chapter applies that argument to a model call. The mapping is ours, not the paper’s:

  • Transport success, a parsed body, a normal finish reason, an answer of the right shape and a clean interpretation are the per-hop checksums — the layers Chapter 12 separated. They are worth having. Chapter 11’s retry decisions and Chapter 12’s truncation detection depend on them.
  • None of them is the task-level check. The end is the task. Evidence used for acceptance has to be bound to the artifact that will be used and evaluate the declared criteria it claims to cover. The acceptance decision then follows from that evidence.

Saltzer and colleagues treated lower-level reliability as useful for performance rather than sufficient for correctness. Carried over, a succeeded call stands to a completed task the way a reliable network stands to a correctly stored file. The analogy stops at the properties the task-level check actually covers: a passing check does not make undeclared properties true or make inadequate criteria adequate.

Which bytes, from where

The end-to-end argument says where the check must happen. Knowing that the checked thing is the thing you are accepting takes a separate provenance rule.

The W3C PROV data model supplies the vocabulary (Moreau and Missier, 2013):

  • An entity is a thing with some fixed aspects.
  • Generation is the completion of production of a new entity by an activity.
  • Derivation turns one entity into another.
  • Association assigns an agent responsibility for an activity.
  • Attribution ascribes an entity to an agent.

Put the paragraph repair into those terms. The call and its final attempt are the generating activity. The artifact is an entity, and its fixed aspect is its SHA-256. A check is an activity that used that entity. An edited paragraph is not the same entity with a small change. In PROV’s terms it is a different entity derived from the first. PROV supplies that distinction; what a check on one entity implies about another is decided by the rule built on top of it, which here is CodeAI’s: nothing established about the first entity transfers to the second.

This is why an acceptance in this chapter does not say “the paragraph is fine”. It names an exact chain:

task → call → final attempt → interpretation → artifact sha256 → check(s) → acceptor role

The rule

In CodeAI this becomes one operation and one projection, task-acceptance-v1. Through Runtime.accept_task, an acceptance is appended as task.accepted only if every reference validates against the ledger:

  • The supplied authority object must allow a new capability, accept, separate from the capabilities that produce work.
  • The acceptor’s actor label must differ from the actor label that produced the output.
  • The task must exist, with the acceptance citing the hash of its declared criteria. You accept against the criteria that were set, not against criteria you prefer.
  • The source call belongs to this task, and its adopted status is succeeded.
  • The cited attempt is the call’s final attempt.
  • The cited interpretation is the interpretation named as the basis of the adopted call-status decision, and it records complete generation.
  • The artifact’s SHA-256 equals the output preserved in that attempt’s envelope, and the stored bytes are intact.
  • At least one check is cited. Each was recorded for this task, targeted artifact:sha256:<those bytes>, and completed PASS.

In the experiment’s producer, the check and the acceptance are two ordinary calls. Reduced from the executed code:

check = runtime.run_check(CheckRequest(
    check_id=check_id, task_id=task_id,
    command=(sys.executable, "paragraph_criteria_check.py",
             "--path", artifact_path, "--sha256", sha),
    target=artifact_target(sha)),              # "artifact:sha256:<those bytes>"
    verifier=LocalCommandVerifier())

completion = runtime.accept_task(AcceptanceRequest(
    acceptance_id=str(uuid4()), task_id=task_id, actor_id="reviewer",
    criteria_sha256=criteria_sha256(CRITERIA),
    artifact_sha256=sha,
    source_call_id=call_id, source_attempt_id=attempt_id,
    source_interpretation_id=interpretation_id,
    check_ids=(check_id,)),
    authority=Authority(frozenset({Capability.ACCEPT})))

Every field in the acceptance is a reference or identity the runtime can look up. Nothing in the request establishes that the paragraph is good in general. The cited checks supply evidence about declared properties; the acceptance is the decision that cites that evidence and binds it to this task, call and artifact. The runtime validates those references, bindings and verdicts. It does not establish properties the checks never tested.

If all of that holds, the runtime appends task.accepted and then task.completed, whose causation_id points at the acceptance. If a submitted acceptance fails validation, it appends task.acceptance_rejected with the reasons; it does not append a new acceptance or completion for that request.

Completion is then not a field anyone sets. It is derived:

Kill it, reopen it: still not done

The demonstration runs offline. The model response is scripted and sent through the real recorded path: the OpenCode adapter, the Chat Completions codec, and the attempt and interpretation machinery from Chapters 11 and 12. That is deliberate. The question here is not whether a model can repair a paragraph. It is what the process may conclude once one has.

Every step runs as a separate operating-system process over one SQLite ledger:

  1. produce creates the task, makes the call, stores the output as an artifact, and writes a checkpoint file only after the ledger has committed the ten events of the task and its call.
  2. The parent waits for the checkpoint. Then the run splits in two:
    • clean exit: the child returns normally, exit code 0.
    • abrupt termination: the child is still running and the parent kills it (TerminateProcess, no cleanup), exit code 1.
  3. inspect is a new process that reopens the ledger and projects the task. It has no transport function to call.

Both runs reopened identically:

clean exitkilled process
Call status / generationsucceeded / completesucceeded / complete
Taskincompleteincomplete
Ledger events1010
Events appended by inspecting00
Model calls made by inspecting00

Reading the process’s state changed nothing and asked the model nothing. Restarting invented no acceptance, and it did not re-run the cognition to be safe. The call’s manifest, start, observation and completion each appear exactly once before and after.

The two interruptions are labeled separately on purpose. The killed process died at a durable checkpoint, after the ledger had committed. That shows the facts survive a process that never shuts down cleanly. It stops short of showing that every crash point is safe, and this chapter makes no such claim.

Evidence, then a decision

A third process verifies and accepts. It runs paragraph_criteria_check.py as its own subprocess, and that script imports nothing from CodeAI. It reads the artifact file, recomputes its SHA-256, refuses the file if the hash differs from the one it was given, and evaluates the three criteria:

artifact_sha256_observed  89de30226cb4b905…  bytes_match: true
no-percentage   passed
marker-once     passed
sentence-kept   passed

A passing check is still not completion. In a separate ledger, one negative case ran the same call and the same passing check and stopped there. It projects incomplete, basis no acceptance recorded. Evidence about an artifact is not a decision about a task.

Then the third process accepts, as a role labeled reviewer holding only accept:

+ task.accepted    actor=reviewer  authority_basis=[accept]  check_ids=[check-293e…]
+ task.completed   actor=runtime   caused_by=task.accepted

Two events. Then it does two things a real system does by accident, and a fourth process looks at the result:

  • Sends the identical acceptance again, with a new request id: 0 events. The projection returned is the same.
  • Runs a second passing check and submits an acceptance citing both checks: refused, conflicting_acceptance. One rejection event is recorded, and there is still exactly one completion.
  • A fourth process reopens: completed, 17 events. The acceptance names the same artifact hash and the same call. The interpretation and the original decisions are unchanged from before acceptance.

The repeat is the part distributed-systems people will recognize. Helland’s position paper on building without distributed transactions starts from the assumption that messages arrive at least once. A recipient must be designed to ignore redundant messages. It typically does that by remembering what it has already processed and answering a repeat the way it answered the original (Helland, 2007). CodeAI applies that reasoning to acceptance. What it remembers is the acceptance’s identity: everything about it except the request id. A reviewer who clicks twice, or a script that retries after a timeout, gets the decision already on record, not a second one.

How CodeAI handles the gap between the two appends is its own design, tested in this experiment rather than taken from the paper. task.accepted and task.completed are two writes, and this runtime does not pretend they are one. A negative case injected a failure after the acceptance was written and before the completion. On reopen, the task projects incomplete, with acceptance_pending_completion: true. Sending the identical acceptance again appended the one missing task.completed and nothing else.

That is a recoverable protocol, not an exactly-once transaction. The failure was an injected exception, not a killed process, and the ledger has one writer. Chapter 16 takes up restarting versus repeating properly.

Twelve ways to cheat

Each case below ran in its own ledger. Apart from the last row, none produced an acceptance, and none projects as completed:

AttemptWhat the runtime recordedComplete?
Succeeded call and passing check, nobody acceptsnothing to reject; projection has no acceptanceNo
Acceptance cites no checkmissing_checkNo
Output kept “73%”; the real check returned FAILcheck_not_passed:…:FAILNo
The check command could not runcheck_not_passed:…:ERRORNo
Passing check on the right bytes, recorded for a different taskcheck_wrong_taskNo
Acceptance cites a different task’s call and outputsource_call_wrong_taskNo
An edited paragraph that passes all three criteriaartifact_not_source_outputNo
Acceptor holds every capability except acceptunauthorizedNo
The producing actor’s label tries to acceptself_acceptanceNo
Generation truncated (finish_reason=length), text still passes the checksource_call_not_succeeded, generation_not_completeNo
A bare task.completed appended by handignored: no causing acceptanceNo
Failure between acceptance and completionpending; identical repeat repairs itOnly after the repeat

Three of these carry most of the argument.

A succeeded call is incomplete. No failure is needed to show this: the call succeeded and the work was not done. Any pipeline that turns a model’s success into a task’s success has merged two facts the ledger keeps apart.

The edited artifact passed its check and was still rejected. The edit was small: “is intended to make” became “should make”. The criteria check passed on the edited bytes, and that check was perfectly valid. It was also about a different entity from the one the call generated. Verification of some bytes is not evidence about different bytes. Chains break exactly here in practice: someone tidies the output after review, or a formatter runs after the tests, and a green check is attached to something it never saw.

A bare task.completed does not complete anything. Anyone with write access to a status field can set it to done. Here someone appended the completion event directly, and the projection asked what acceptance caused it. There was no answer, so the task stayed incomplete. The ledger is not replaying commands. It interprets evidence according to a protocol.

The truncated case deserves a line too. The text happened to pass the criteria. The generation still did not finish, and the runtime refused to accept output from an unfinished generation just because what arrived looks acceptable.

A separate verifier that shares no code with CodeAI reconstructed the same completion chains from the raw ledgers; the evidence notes say what it checked.

Completion is not something the model reports. It is state the process derives from its recorded history under an explicit protocol and the trust assumptions of that record.

Roles, not yet authority

Be precise about what the accept capability and the self-acceptance refusal establish.

The runtime now represents producer and acceptor as distinct roles, and it refuses to let one actor label fill both. That is structure, and it is useful.

It is not security:

  • A label is a string, not an authenticated identity.
  • Ledger events are unsigned. Anyone who can write the SQLite file directly could append a matching task.accepted and task.completed pair, and the projection would believe it. The ledger is the trust boundary.
  • The criteria check was written by the same author as the task, runs on the same machine, and checks mechanical criteria. It is a deterministic procedure, not an independent judge of whether the paragraph is true. Chapter 21 splits that worry into three questions — is a check independent of what it evaluates, adequate to the property that matters, and bound to the exact state being accepted? — and this check answers them unevenly: bound to the bytes by hash, adequate only to three mechanical rules, independent of the model but not of the task’s author.

Who is actually permitted to make a consequential decision is Chapter 20’s subject. What makes evidence independent of the thing it evaluates is Chapter 21’s. This chapter needed only one step: that “done” is a decision with recorded inputs, not something a component announces.

Where it is still weak

  1. Labels are not identities. reviewer and repairer are strings chosen by the caller.
  2. The ledger is unsigned. A forged acceptance–completion pair would project as complete.
  3. Single writer only. Nothing prevents two processes accepting at the same moment.
  4. Crash coverage is narrow. The kill happened at a durable checkpoint, and the window between the two appends was exercised by an injected exception, not a kill.
  5. The check is mechanical. It shows the bytes meet three written rules. It says nothing about whether the claim is now true, or whether the rules were the right ones.
  6. The model output was scripted. The recorded path is real. No live model was asked, so nothing here measures repair quality.
  7. Policy v1 accepts only the call’s own output. A human-edited paragraph cannot be accepted under it. The next version would need to record the derivation and who made it.
  8. One task, one artifact, one kind of check. There is no multi-artifact completion, no partial completion, no run status (list_runs still reports active), and no CLI surface.

Do this now

Forty minutes. Find out what “done” means in your system.

  1. Take the last ten items your team, pipeline or agent marked done. For each one, write down the exact version of the artifact, the check that ran against those bytes, and who decided. Count the cells you cannot fill.
  2. Search for the point where a model call’s success, an HTTP 200 or a green CI run becomes a status called done, complete or resolved. Note whether anything records what that status was based on.
  3. In a disposable environment, change the artifact after its check passes. Does anything notice?
  4. Set the done status directly, without doing the work. Does anything ask why?

If you are building with an assistant:

Add explicit task acceptance on top of an append-only event log.
A successful model call or a passing check must never complete a task.
- Acceptance names: task id, hash of the task's declared criteria, exact
  artifact sha256, source call and final attempt, the interpretation the
  call's status decision used, and one or more check ids.
- Validate against the log before recording anything: same task; call
  succeeded; generation complete; artifact equals the call's preserved
  output and is intact; each check was for this task, targeted these bytes
  and passed; the acceptor holds an accept permission and is not the
  producer.
- On success append "accepted", then "completed" with a causal link to it.
  On failure append "acceptance rejected" with reasons and nothing else.
- Derive completion only from a completed event caused by an acceptance.
  A completed event with no cause does not count.
- Identical repeats append nothing and repair a missing completion;
  a different acceptance for an accepted task is refused.
- Test in separate processes: produce, kill, reopen (still incomplete, no
  model call), check, accept, reopen (complete). Then try each negative case.
- State the limits: labels are not authentication; single writer; no
  exactly-once guarantee.

Failure modes

  • Call success as task success. The call succeeded; the work had not been decided.
  • A passing check as completion. Evidence about an artifact is not a decision about a task.
  • Checking different bytes than you accept. The edited paragraph passed its check.
  • A status field anyone can set. Completion without a recorded cause is a claim, not a fact.
  • Letting the producer accept its own output. Even as labels, keep the roles apart.
  • Accepting against criteria you changed afterwards. Bind acceptance to the criteria as declared.
  • Non-idempotent acceptance. A double click produces two completions.
  • Pretending two writes are one. Make the gap recoverable and visible.
  • Accepting truncated output because it looks fine. Unfinished generation is not a finished artifact.
  • Calling a label an identity. Roles are structure, not security.

What this chapter established

  • A successful call is not finished work. Generation succeeded ≠ artifact checked ≠ artifact accepted ≠ task completed. The output belongs to the call; “done” belongs to the task.
  • Check at the end. Transport success, a normal finish and a well-formed answer are lower-layer reliability. The task-level check belongs at the task boundary, bound to the artifact that will actually be used and the declared criteria it covers. A pass establishes those checked properties, not general correctness. That is Saltzer, Reed and Clark’s end-to-end argument, applied here by the book.
  • Check the bytes you accept. An acceptance names the generating call, the exact artifact hash, the checks run on those bytes and an eligible role. An edited artifact is a different entity, and a check on one says nothing about the other.
  • Derive “done”; do not set it. Completion counts only when it was caused by a recorded acceptance. A completion with no cause is a claim. An identical repeat must change nothing, and the gap between accepting and completing must be recoverable, because the two are separate writes.
  • Roles are structure, not security. Keeping producer and acceptor apart matters, but a label is not an authenticated identity and an unsigned ledger is a trust boundary.

What CodeAI showed. Acceptance was built as task-acceptance-v1. In separate processes, with both a clean exit and a killed producer, reopening showed a succeeded call and an incomplete task; inspecting appended 0 events and made 0 model calls. A check plus an acceptance appended 2 events and completed the task. An identical repeat appended nothing, a conflicting acceptance was refused, and history before acceptance was unchanged. Twelve negative cases, including an edited artifact with its own passing check and a hand-appended task.completed, did not complete a task. The model output was scripted, so nothing here measures repair quality, and authentication, signing, concurrency and crash atomicity are not established.

Evidence notes

Separate reconstruction. A verifier sharing no code with CodeAI opened every ledger read-only, with the stored bytes, and reconstructed the chain:

  • one acceptance and one completion in each run, the completion caused by and sequenced after the acceptance
  • the accepted hash equal to the candidate bytes, which equal the text in the stored transport body
  • every cited check linked, recorded for the task, aimed at those bytes, passing, and earlier than the acceptance
  • history before acceptance an unchanged prefix of history after
  • across the negative ledgers, exactly one causal acceptance–completion pair: the repaired interruption

It reached the same completion result as Runtime.task_completion(). This is an implementation cross-check of the recorded causal chain, not evidence that the ledger is authentic or that the task’s criteria are adequate to every property that matters.

Next

Part 2 took the model out of the chat box one boundary at a time. Chapter 11 recorded the call. Chapter 12 survived several dialects. Chapter 13 made the numbers honest. This chapter put the call inside a process that can say, from its own history, whether the work is done.

Look again at what the acceptance cites: a call, an attempt, an interpretation, an artifact, checks. It does not cite what the model was shown. The check establishes that the output meets the criteria. Nothing in the chain establishes that the input was the right input, or lets another process reconstruct it. Part 3 starts there.

Continue with What Did the Model Actually See?.

References

Implementation sources: CodeAI call identity repair (Runtime._interpret_attempt; tests/test_opencode_gateway.py::test_recorded_call_completed_carries_logical_call_id) and acceptance feature (src/codeai/acceptance.py: task-acceptance-v1, AcceptanceRequest, TaskCompletion, accept_task, project_task_completion; Capability.ACCEPT in src/codeai/domain.py; Runtime.accept_task and Runtime.task_completion; tests/test_task_acceptance.py; experiments/task_completion_demo.py; experiments/paragraph_criteria_check.py). Suite 289 passed. Evidence: experiments/applied-ai/evidence/task-completion/offline-fe0797d/, run from committed code (manifest, both process chains with checkpoints, termination records, before/after projections and event exports, check reports, acceptance chains, negatives/, hashes.json), and independent-verification/ (a verifier sharing no code with CodeAI, and its result). Stage 12 ledgers showing the empty call id: experiments/applied-ai/evidence/protocol-conformance/live/*/events.json.