← Applied AI

Claims, Evidence, and Decisions

Said is not supported, and supported is not relied on. A claim is attributed to exact preserved bytes and starts unresolved; evidence is recorded under explicit provenance rules without pretending that provenance proves entailment; a decision snapshots the claims it relied on so later changes to that basis remain inspectable.

Part 3 — Give Intelligence a Runtime

What did the decision rest on?

AI output feeds decisions. A review says a setting is correct, a summary says the tests pass, an assistant says a change is safe, and someone merges. When one of those statements turns out to be wrong, asking whether the AI was wrong gets you very little. What you need to know is which statements the decision actually rested on, and what anyone had checked. Most systems cannot tell you, because the decision was recorded without its basis.

Three different things hide inside the word “claim”:

said  ≠  supported  ≠  relied on

A statement that was said, even quoted exactly, is only attributed. In CodeAI, support requires a separate evidence record: a passage in a different preserved source, or a completed check whose request named the claim. The runtime can validate those provenance and binding facts. It does not establish that the passage entails the claim or that the check is adequate.

By the end of this chapter you will be able to record AI-derived claims so that each one points at exact preserved text, evidence is recorded separately from the claim, and a decision snapshots the claims it explicitly relied on. A decision may also name unresolved claims it knowingly leaves open. That list is caller-supplied, not an exhaustive proof of everything the decision-maker considered. When recorded evidence or a source call changes later, the system can name decisions whose recorded basis moved without rewriting those decisions.

Three sentences and a merge

A review of a cache change says three things:

The cache TTL is set to 60 seconds in config/cache.toml. The retry test passes.
p99 latency stays under 200 ms.

Someone reads it and merges the change. The next day, the deployed configuration turns out to say 600 seconds.

The obvious question is whether the review was wrong. The more useful question is narrower: which part of the decision rested on which sentence, and what had anyone actually checked? One sentence could be confirmed against a file. One could be tested. One had no evidence at all. The merge used two of them and knowingly left the third open, or it didn’t. The merge record lists no basis, so the basis cannot be reconstructed.

Which statements are supported, by what, under which conditions?

Said, supported, relied on

Three things are easy to collapse into one word, “claim”.

Said. Someone, or some model, stated something, at a place you can point to. The W3C provenance model gives this a precise name. In its words, “Attribution is the ascribing of an entity to an agent” (Moreau and Missier, 2013). It separately defines a quotation as the repeat of some or all of an entity. Attribution tells you who said it and where. It tells you nothing about whether it is true.

Supported. Something other than the statement bears on it. FEVER, a large fact-verification dataset, labels each claim Supported, Refuted or NotEnoughInfo (Thorne et al., 2018). For the first two labels, annotators also recorded the sentences that formed the evidence. The third outcome matters as much as the other two: a claim with no evidence is not false, it is unresolved.

FActScore breaks long generations into atomic facts and measures the share supported by a knowledge source (Min et al., 2023). Its central observation is that one fluent answer can hold supported and unsupported statements side by side.

Relied on. A decision used the statement. This is the part the evaluation literature doesn’t cover, because benchmarks score claims and don’t act on them. A system that acts has a further duty: to record which claims it used, in what state, so that a later change to those claims can be traced to the decision.

None of these sources is about model runtimes, and CodeAI does not do what FEVER and FActScore do. Their verdicts come from annotators or from a model estimating support. CodeAI makes no judgment of entailment at all. It checks that evidence exists where it says it does, and who recorded it. The judgment that a passage bears on a claim stays with a named actor. The mapping in this chapter is the book’s.

LayerRecorded asWhat the runtime can checkWhat it cannot check
Saidclaim.extractedThe quote is the preserved response text at its span; the bytes match their hashThat the statement is atomic, or paraphrases the quote faithfully
Supportedclaim.evidence_recordedA passage exists in a different, hash-matched source; a check named the claim and its verdict matches; the recorder did not produce the claimThat the passage entails the claim, or that the check tests it adequately
StandingprojectionStatus and evidence class follow from the evidence, and the source call’s adopted statusTruth
Relied ondecision.recordedThe policy held when the decision was made; a snapshot of what it heldWhether acting on a changed basis was wrong

The vocabulary still runs from E0 to E4, but the Stage 18 path enters at E1. E2 and E3 are evidence classes, not mandatory sequential gates: a targeted completed check can take an attributed claim directly to E3 without a prior E2 source passage. E4 has no construction path here. Standing is derived, never stored.

How CodeAI records claims, evidence and decisions

CodeAI could not answer those questions either, at first. It had the vocabulary: evidence classes from E0 (asserted) to E4 (robust), claim statuses, even a Decision type. A probe of the earlier claim API showed what the vocabulary was worth:

  • Asserted evidence was accepted. A claim recorded with evidence class E4_ROBUST and status “supported”, citing a source artifact that did not exist, was stored and projected exactly as asserted.
  • Any passing command promoted. A claim was promoted to E3_REPRODUCED by a check whose command printed “nothing tested”.
  • No decisions. There was no way to record one at all.

Reading the old projection’s source adds two more, not part of the preserved run. A later claim.recorded for the same ID replaces the claim, refuted or not. And promotion events are accepted from any actor.

The repair adds src/codeai/evidence.py and five runtime operations beside the old claim API, which is left as it was. The new projections ignore the old API’s events and list them, so a reader sees both.

The shape is clearest in the calls the experiment actually made. Reduced from the executed producer, with comments naming the positional arguments:

# Said: an exact span of a preserved response. No evidence class, no status.
runtime.extract_claim(ClaimExtraction(
    "c-ttl", TASK, call_id, attempt_id,
    0, 56, quote,                          # span start, span end, quoted text
    statement, "claim-extractor"))

# Supported: a passage of a different preserved source, cited by a named actor...
runtime.record_claim_evidence(EvidenceRecord(
    "ev-config", "c-ttl", SOURCE_PASSAGE, SUPPORTS, "human-reviewer",
    source_artifact_id=config.artifact_id,
    passage_start=start, passage_end=start + len("ttl_seconds = 60"),
    passage="ttl_seconds = 60"))

# ...or a completed check whose request named the claim.
runtime.record_claim_evidence(EvidenceRecord(
    "ev-retry", "c-retry", CHECK, SUPPORTS, "verification-reviewer",
    check_id="check-retry"))

# Relied on: the claims a decision uses, and the ones it knowingly leaves open.
runtime.record_decision(DecisionRequest(
    "dec-merge", TASK, "release-manager", "Merge the cache change",
    ("c-ttl", "c-retry"),                  # relied on
    ("c-latency",)))                       # acknowledged open

runtime.decision_standing("dec-merge")     # basis_intact or basis_changed, naming what moved
runtime.decisions_resting_on("c-ttl")      # every decision that used this claim

extract_claim reads the response bytes the runtime preserved — the same bytes Chapter 17 insists must outlive every interpretation — checks their hash, and refuses unless the quote equals the output text at that span. The claim is causally linked to the observation and starts unresolved at E1_ATTRIBUTED, whatever it says. The caller cannot supply an evidence class or a status.

record_claim_evidence accepts two kinds of evidence. A source passage (E2_SOURCE_CHECKED) must be an exact passage of a preserved, hash-matched artifact that is not the claim’s own response: a review cannot cite itself. A check (E3_REPRODUCED) must be a completed check whose request named the claim; PASS may support, FAIL may refute, and other verdicts are refused as evidence. The evidence recorder’s actor ID may not be one of the actors recorded as producing the claim’s source call. These are provenance and binding checks. For a source passage, whether it bears on the claim remains the recorder’s judgment; for a command check, naming the claim does not establish that the command tested it adequately.

Standing is derived, never stored. claim_standing reads status from the evidence — supports only is supported, refutes only is refuted, both is contested, neither is unresolved — takes the strongest supporting evidence class, and adds the source call’s adopted status, which Chapter 17 made revisable. A refutation changes standing without pretending that the runtime has adjudicated between conflicting sources.

record_decision refuses unless every claim listed in relied_on_claim_ids is supported, at E2 or better, from a source call whose adopted status is succeeded. The separate acknowledged_unresolved_claim_ids list is optional caller-supplied context: the runtime refuses overlap with relied-on claims, but it does not prove that the list is exhaustive or even require each acknowledged claim to be unresolved. The decision stores a snapshot of each relied-on claim’s standing, and decision_standing and decisions_resting_on compare those snapshots with current standing. They report basis_intact or basis_changed, name every tracked field that moved, and never edit or revoke the decision.

Every refusal is appended as its own event — claim.refused, claim.evidence_refused, decision.refused — and raised. Nothing else is appended.

One decision, followed over two days

The run was preregistered, then executed. Every step — review, extract, evidence, decide, contradict, reinterpret, inspect — ran as its own operating-system process, so nothing survived between steps except what the ledger and artifact store held. A synthetic provider logged every request to a receipt log outside the ledger, outbound connections were refused, and the checks were real local commands: parsing the TOML file and a small retry test. The review itself is a synthetic fixture.

CaseWhat happenedResult
Decision over daysReview, extract, evidence, decide; day 2 brings a contradicting sourceDecision recorded with its basis; day 2 changes its standing on one claim; the record is unchanged
AgreementThe same sentence from two different modelsBoth unresolved; the decision relying on the agreement is refused
Source reinterpretedThe review first recorded as succeeded under Chapter 17’s v1 interpreterAfter reinterpretation, the decision’s standing changes on its source; an identical new decision is refused
Bytes deleted, bytes corruptedThe response body damaged before extractionEvery extraction refused
Before Stage 18The old API from Stage 17, then Stage 18 on the same ledgerAsserted E4 and blanket E3 accepted before; a decision citing them refused after

Claims that start unresolved

The review call produced one provider receipt and ten events, and a separate process then extracted three claims — characters 0–56 (the TTL), 57–79 (the retry test) and 80–111 (latency) — all unresolved at E1_ATTRIBUTED.

The same process tried to extract a fourth claim, quoting the review as saying “600 seconds”. The quote did not match the preserved text at that span, so the attempt was refused for quote_not_in_source, with only the refusal in the ledger and no claim created. Repeating a real extraction appended nothing, and with no further provider requests after the review, the receipt count stayed at one through every later step.

Evidence that counts, and evidence that doesn’t

Counted, from a source passage. A human reviewer cited the passage ttl_seconds = 60 in the preserved config/cache.toml. The TTL claim became supported at E2_SOURCE_CHECKED.

Three attempts were refused, each appending only its refusal:

  • Self-evidence. The reviewing model recorded the same passage as evidence for its own claim: self_evidence.
  • Circular evidence. The human reviewer cited the review’s own response bytes as the source for the TTL sentence. The passage was really there, but a statement does not become evidence for itself by being quoted: source_is_claim_origin.
  • An untargeted check. A check that parsed the TOML file, naming only the TTL claim, passed. Cited as evidence for the retry claim, it was refused: check_not_targeting_claim. Before this stage, a passing check was enough.

Counted, from a check. The retry test ran (exit 0, “retry test passed”) under a check request that named the retry claim. Cited for that claim, it made the claim supported at E3_REPRODUCED.

run_check still appends its old promotion events as it always did. The new standing ignores them and lists their IDs, so the discrepancy is visible rather than silent.

A decision with its reasons attached

A release manager first tried to ship on the TTL and latency claims. Refused: claim_not_supported:c-latency and evidence_below_policy:c-latency. The refusal names the claim and records what the manager was trying to rely on.

The second decision, “Merge the cache change”, relied on the TTL claim (supported, E2) and the retry claim (supported, E3). It acknowledged the latency claim as unresolved. It was recorded with a snapshot of both relied-on claims:

  • status and evidence class
  • the evidence IDs for and against
  • the quote and observation hashes
  • the source call’s adopted status and the record it came from

Repeating it appended nothing.

That acknowledgement makes one omission explicit: for this decision, latency was recorded as open rather than silently treated as support. It does not prove that every unresolved consideration was listed. CodeAI can enforce the basis the caller declares; it cannot infer an undeclared reason that influenced the human decision-maker.

Day two: the source disagrees

A new process recorded one more piece of evidence. An on-call engineer cited ttl_seconds = 600 in the deployed configuration, a different preserved file, as refuting the TTL claim. The claim became contested. It still has its E2 support; it now has a refutation too, and CodeAI does not choose between them.

The decision’s standing became basis_changed, with exactly two changes on the TTL claim — status: supported → contested, and refuting_evidence_ids: none → ev-deployed — while the retry claim’s entry did not change.

Asked which decisions rest on the TTL claim, the runtime names the merge. The retry claim names it too without a change to its entry, and the latency claim names none, because the merge acknowledged it and never relied on it.

The decision.recorded event is byte-for-byte the event recorded the day before, and so are all 27 events before it. The decision is not rewritten to fit what is now known. It is shown to have been made on a basis that has since moved, and exactly where.

A moved basis is not the same as a defeated one

Reporting that a basis moved is the projection’s job. Deciding what a particular move means is a separate question, and the difference matters as soon as anything acts on a decision. basis_changed fires when any tracked field differs — including when a claim gains more support, which is a decision becoming better founded rather than worse.

CodeAI later separated the two, because an effectful operation that cites a decision needs a verdict rather than a diff:

Since the decisionThe basis is
New supporting evidence, or a higher evidence classintact
Refuting evidence recorded afterwardsdefeated
The claim now stands refuted or contesteddefeated
Evidence class fell below what was recorded, or below E2defeated
The source call’s adopted status is no longer succeededdefeated
The claim can no longer be projectedunknown, which is also a refusal
A verification attempt returned INCONCLUSIVE or ERRORnothing — an attempt is not standing

That last row is the one to hold onto. A check that could not run, or could not settle the question, moves no claim; letting it revoke a decision would hand an infrastructure failure the power to withdraw a justification. The runtime keeps attempt, evidence and standing apart precisely so that this line can be drawn (Chapter 21).

New evidence can defeat a decision. It cannot silently rehabilitate one. If the TTL claim were later supported again by better evidence, the merge decision would stay defeated: it was a judgment made against an evidentiary state that has since been overturned, and reviving it because the current projection happens to look positive again would be exactly the retroactive rewrite this chapter refuses. The repair is a new decision, not a resurrection.

And that is where this chapter’s model currently runs out. A claim carrying both support and refutation stands contested, and record_decision will not rest a new decision on a contested claim either. So the refutation does not merely retire the merge decision — it blocks any decision resting on that claim, and nothing in the runtime settles a contest. A contested claim needs an act of adjudication, with its own authority and provenance, before it can justify anything again. That mechanism does not exist, and inventing it as a side effect of some other seam would be the wrong way to get it.

When the source call is reinterpreted

Chapter 17 showed that a call’s adopted status can change after the fact. This run joins the two chapters.

  1. Day one. The review was recorded as succeeded under the old interpreter, although its finish reason was length.
  2. Claims, evidence, decision. Claims were extracted, both were supported as before, and the merge decision was recorded.
  3. Reinterpretation. reinterpret_call then read the preserved bytes under v2 and adopted unresolved, without a provider request.

The decision’s standing changed on source_call_status, succeeded → unresolved, for both claims, along with the status record each now rests on. The evidence had not changed. The source had.

An identical decision recorded afterwards was refused for source_call_not_succeeded on both claims.

Agreement is not evidence

Two different models, on two separate calls, returned the same TTL sentence. Both claims were attributed and both stayed unresolved at E1. A decision relying on “two models agree” was refused with four reasons: neither claim was supported, and neither met the evidence policy.

Two quotations of an unchecked statement are still two unchecked statements. Agreement can tell you what to check next. It is not the check.

Without the bytes

Two copies were taken right after the review: one with the response body deleted, one with its bytes altered. All five extraction attempts on each were refused for observation_unavailable, and no claim was created.

A claim that points at text nobody can read back is not attributed. It is only asserted.

What this is not

  • Fact-checking is out of scope. CodeAI confirms that a passage exists in a different, preserved source and who cited it. Whether the passage entails the claim is that person’s judgment, recorded under their name.
  • Not claim extraction. Spans were supplied. Nothing here splits text into atomic facts, and the statement’s paraphrase is not checked, only the quote.
  • Stage 18 is reporting, not action enforcement. The pinned run reports basis_changed and does not block an effect. Later work adds a narrower decision-evidence gate: an action that explicitly cites a decision can be refused when that recorded basis is defeated or unknown. Actions that cite no decision remain outside that gate, and nothing is revoked retroactively.
  • Not truth maintenance. Evidence is never retracted, and a newer source does not outrank an older one. Two disagreeing sources simply make a claim contested.
  • Not a fix of the old API. record_claim still stores whatever it is told, and experiments still use it. Decisions refuse to rely on its claims.

Where it is still weak

  1. The self-evidence rule compares actor IDs. The same model or person under a different ID passes.
  2. A check supports a claim because the claim was named in its request. Whether the command tests the claim is not examined.
  3. The decision policy is fixed at E2, and E4_ROBUST is unreachable.
  4. Spans are character offsets into text derived from the bytes by CodeAI’s output reader, not byte offsets into the raw response.
  5. Contested claims have no adjudication. Support plus refutation is a stable end state: no new decision may rest on such a claim, and the runtime offers no operation that settles the contest. The projection is honest about the standoff and cannot resolve it.
  6. One synthetic review, local checks, a single writer.

Do this now

Thirty minutes. Take one decision made in the last week on the strength of an AI answer.

  1. List the statements the decision actually used. For each one, point to where it was said: a message, a line, a file. If you can’t point to it, it is asserted, not attributed.
  2. For each statement, write what supported it: a document passage you read, a test you ran, or nothing. Mark anything supported only by “the model said so” or “two tools agreed” as unresolved.
  3. Mark the statements the decision knowingly left unresolved. Were any of them relied on anyway?
  4. Suppose one supporting source changed tomorrow. Could you find this decision from that source, without remembering it?

If you are building with an assistant:

Separate what was said, what supports it, and what a decision relied on.
- Record a claim only as an exact quote at a span of preserved output,
  hash-checked; it starts unresolved. Refuse quotes not at the span.
- Record evidence as its own event: a passage of a different, preserved,
  hash-matched source, or a completed check whose request named the claim.
  Refuse evidence recorded by the claim's producer, evidence citing the
  claim's own output, and checks that did not target the claim.
- Derive status (supported/refuted/contested/unresolved) and evidence
  class from evidence records. Never store an asserted status.
- Record a decision with the claims it relies on and those it leaves
  open; refuse reliance on unsupported claims; snapshot their standing.
- Add a projection that compares each decision's snapshot with current
  standing and names what changed. Never edit the decision.
- Test in separate processes: refusals appended and nothing else; a later
  refutation changes the decision's standing and not its record.

Failure modes

  • Treating attribution as support. It records who said it, not whether it holds.
  • Letting agreement stand in for evidence. Two unchecked quotations are still unchecked.
  • Letting any passing check promote. The check must be about the claim.
  • Citing the output as its own evidence. A quote of the claim is not a source for it.
  • Storing a status instead of deriving it. Whoever writes last decides what is true.
  • Deciding without recording the basis. When a source changes, no one can find what rested on it.
  • Rewriting the decision when the basis moves. Report the change; keep what was decided and why.

What this chapter established

  • Said ≠ supported ≠ relied on. Attribution records who said something and where (PROV-DM). CodeAI records separate evidence and derives standing from it; the runtime does not establish entailment or check adequacy. “Not enough information” remains an honest claim-level outcome in the literature (FEVER, FActScore). The runtime mapping is the book’s.
  • Attribute to exact text. A claim points at an exact span of preserved response text and starts unresolved, whatever it says. A claim whose source bytes cannot be read back is not attributed by this path.
  • Evidence is separate and provenance-checked. A passage in a different preserved source, or a completed PASS/FAIL check whose request named the claim, can be recorded. The claim producer cannot record its evidence under the same actor identity; the claim’s own response cannot serve as its source passage; an untargeted check and agreement between two models do not become evidence merely by existing.
  • Derive standing; snapshot declared reliance. Status comes from the evidence records, not from whoever wrote last. A decision snapshots the claims it explicitly relied on. It may also record open claims the caller chooses to acknowledge, but that list is not an exhaustive account of everything the decision-maker considered.
  • Keep the decision when the basis moves. project_decision_standing reports any tracked difference without rewriting the decision. Later project_decision_evidence distinguishes benign added support from a defeated or unknown basis for future reliance; neither mechanism rewrites past decisions or effects.

What CodeAI showed. In the Stage 18 pinned run, separate processes refused and recorded a forged quote, self-authored evidence, circular evidence, an untargeted check, agreement between two models as a decision basis, missing or corrupt response bytes, and a decision relying on an unresolved claim. A recorded merge decision’s standing changed when a deployed configuration refuted one claim, and separately when its source call was reinterpreted; the decision record stayed byte-for-byte unchanged and named exactly what moved.

That run did not establish entailment, automatic claim extraction, action enforcement, source precedence, or a repair of the legacy claim API. Later W2-3 work added the narrower defeat projection and a gate for actions that explicitly cite a decision; it still does not require every action to cite one, retroactively undo effects, or settle contested claims.

Evidence notes

Independent verification. The bundle’s verifier imports neither CodeAI nor the producer and applies its own rules: every quote checked at its span in the preserved response, every evidence item validated against source bytes or check records, and every claim standing, decision basis, basis change and refusal reason re-derived from the events exported at that moment. It also requires each later export to extend the earlier one unchanged, with provider receipts coming only from review steps.

All 17 semantic claims pass; the full run, including byte hashes for every file in the bundle, exits 0. A fresh CodeAI process reopened copies of three ledgers and reproduced every recorded standing.

Five seeded corruptions were run with the byte inventory bypassed:

CorruptionClaims that failed
The TTL quote forged to “600 seconds”attribution; decision basis; history
The config evidence re-attributed to the reviewing modelevidence validity; history
The latency claim’s recorded standing inflated to supported, E3projections
The decision’s basis rewritten to show the TTL claim already contestedprojections; decision basis; day-2 change; history
The retry check’s request stripped of the claim it namedevidence validity; history

Next

A decision can now say what recorded claims it rested on. It still changes nothing. Merging, deploying and writing a file are effects, and Chapter 16 showed why a process must treat external effects separately from the conclusions that motivated them.

The next construction step is narrower than the later decision-evidence gate described above: record an action request, keep the worker’s report separate from an observation of resulting state, and ask what actually changed. At Chapter 19’s stage, a decision reference is only an application convention in the action payload; the runtime does not yet enforce a decision-to-action link, and explicit authority is developed in Chapter 20.

Continue with The Agent Said Done. Did Anything Change?.

References

  • Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. EMNLP, 2023. arXiv:2305.14251. https://arxiv.org/abs/2305.14251
  • Luc Moreau and Paolo Missier (eds.). PROV-DM: The PROV Data Model. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-dm/
  • James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. FEVER: a large-scale dataset for Fact Extraction and VERification. NAACL-HLT, 2018. arXiv:1803.05355. https://arxiv.org/abs/1803.05355

Implementation sources:

  • CodeAI:
    • src/codeai/evidence.py: ClaimExtraction, EvidenceRecord, DecisionRequest, ClaimStanding, DecisionStanding, extract_claim, record_evidence, project_claim_standing, record_decision, project_decision_standing, decisions_resting_on, CLAIM_EVIDENCE_V1, DECISION_BASIS_V1.
    • src/codeai/runtime.py: extract_claim, record_claim_evidence, claim_standing, record_decision, decision_standing, decisions_resting_on, decision_evidence.
    • Defeat semantics (later work): src/codeai/evidence.py: project_decision_evidence, DecisionEvidenceStanding, BasisClaimStanding, DECISION_EVIDENCE_V1; tests tests/test_decision_evidence.py (14); design and results experiments/W2-3-prereg.md, experiments/W2-3-decision-evidence-binding.md, falsifier experiments/W2-3-falsifier-results.json.
      • Stage 18 evidence-run tests: tests/test_claim_evidence.py (14); the then-full suite reported 341 passed. Later W2-3 defeat work reports its own test and full-suite results in experiments/W2-3-decision-evidence-binding.md.
  • Evidence: experiments/applied-ai/evidence/claims-evidence/2026-09-14-3b6d8fb/.
    • Preregistration and execution record.
    • Six cases, each with an SQLite ledger, artifact store, provider receipt log, workspace, per-step action records, inspections and exported events.
    • The independent verify.py, five seeded corruptions, test outputs, chapter-evidence-report.md and hashes.json.
  • Producer: experiments/applied-ai/claims_evidence_demo.py. The executed copy is pinned as the bundle’s run.py.