← Applied AI

What Did the Model Actually See?

Most systems cannot say what a model was given. Treat context as a compiled input: available, eligible, selected, rendered and received are different states, exclusion is evidence, and a record of what was selected means nothing until it is tied to the bytes actually sent. Tested on CodeAI's context compiler, where the record and the request turned out to disagree.

Part 3 — Give Intelligence a Runtime

A component with a contract still needs a world around it. This part builds that world out of the things that used to live in a person’s head and in the scrollback.

Each chapter makes one of them explicit and durable. What the model was allowed to see becomes a compiled, recorded input rather than an accumulated transcript. Working state becomes something another process can resume from instead of restarting blind. The provider’s response is preserved before anything interprets it, so a corrected reading can be applied later without asking the model again. Statements become claims that point at evidence, and decisions record what they relied on. Then the process is allowed to change something outside itself — and to observe the result rather than trust a report of it.

This is where AI work stops disappearing into scrollback.


Context is an input, not a transcript

When an AI answer is wrong, the first question worth asking is what the model was given. Most systems cannot answer it. They send “the conversation so far”, everything that happened in the order it happened, or they send a prompt string someone assembled from whatever was in scope and nobody recorded. Neither can be checked afterwards, and neither was decided on purpose.

In a chat window that is tolerable, because a person can scroll up. In a process it is a liability. You cannot debug a call, audit a decision, compare two runs, or keep one call from seeing another call’s work if you cannot say what went into the request.

So stop treating context as history. Treat it as an input to be compiled. On its way to a model, a piece of material passes through five states, and each one is a separate question:

available  ≠  eligible  ≠  selected  ≠  rendered  ≠  received

By the end of this chapter you will be able to design context as a deliberate, reproducible input: required material first, every exclusion recorded with its reason, a budget that fails loudly instead of trimming silently, and a record that ties what was selected to the bytes that were actually sent.

The stakes are not hypothetical. When this chapter tested CodeAI’s context compiler, the record said the source paragraph had been selected for the model. The request that would have been sent did not contain it. Nothing failed. The record described an input that never existed.

A transcript is an order you did not choose

A transcript is a context policy by default. It includes everything, in arrival order, until something runs out of room.

Two results suggest why that is a poor policy, even before cost enters.

Liu and colleagues placed the document containing an answer at different positions in long inputs and measured how well models used it (Liu et al., 2024). Performance was often highest when the relevant information came at the beginning or the end, and degraded when it sat in the middle, including for models built for long contexts. Where material lands is part of the input.

Shi and colleagues added irrelevant sentences to grade-school arithmetic problems and built a benchmark from them, GSM-IC (Shi et al., 2023). Accuracy fell markedly when the irrelevant information was present. Sampling several answers, and telling the model to ignore irrelevant information, were among the mitigations that helped. What you include is part of the input too, including what you should have left out.

Both findings are bounded by the models, tasks and prompts they tested. This chapter runs no model and measures neither effect. They are the motivation, not the evidence: if position and irrelevant material change what a model does, then which material went in, and in what order, is a decision worth making deliberately and recording.

Five states of a piece of material

Chapter 14’s acceptance names the call, its final attempt, the exact artifact and the checks run on it. It does not name what the model was shown. Start there.

Take any piece of material a model might use: an objective, a source paragraph, an earlier observation, a claim derived from another call. Before a model can use it, five different questions have to be answered, and they have different answers:

CodeAI’s context compiler

CodeAI’s ContextCompiler already existed. It was not rebuilt for this chapter. What it does, read from the source as it stood for this experiment:

  • Candidates, not history. Each piece of material is a ContextCandidate with a kind (event, artifact, claim), an ID, a declared size, a required flag and the lineage IDs it derives from.
  • A fixed order. Candidates are sorted by kind and then ID. The order they arrive in does not matter.
  • Required first. A required candidate that was not supplied, or that a seal forbids, raises RequiredContextMissing. Required candidates whose declared sizes exceed the budget raise ContextBudgetUnsatisfiable. Neither failure produces a package.
  • Seals by lineage. A Seal lists forbidden event, artifact, call and lineage IDs. An optional candidate whose own ID or declared lineage matches is excluded, with the reason.
  • Budget for the rest. Optional candidates are admitted in the fixed order while they fit; the rest are excluded “because budget”.
  • Two identities. The package ID is a SHA-256 over the task, actor, prompt and version, objective, budget, seal and the sorted selected IDs. The trace ID is a SHA-256 over the package ID and every decision with its reason.

In the Stage 15 runtime exercised here, Runtime.compile_and_record_context compiled first and appended one context.compiled event only on success. Chapter 16 later changes this path to append context.compilation_requested with the offered candidate inventory before compiling, followed by either context.compiled or context.compilation_failed. Keep the stage boundary explicit: the experiment below measures the earlier path.

The pipeline, as a flow of decisions:

    flowchart TD
    CAND["candidates<br/><i>event · artifact · claim, with size + lineage</i>"] --> REQ{"required?"}
    REQ -->|"yes"| RQ["required check<br/><i>missing or sealed → raise,<br/>no package</i>"]
    REQ -->|"optional"| SEAL{"seal forbids<br/>id or lineage?"}
    SEAL -->|"yes"| EX["excluded<br/><i>with the reason recorded</i>"]
    SEAL -->|"no"| BUD{"fits the budget?"}
    BUD -->|"no"| EX
    BUD -->|"yes"| SEL["selected<br/><i>fixed order: kind, then id</i>"]
    RQ --> SEL
    SEL --> PKG["context package<br/><i>id = hash of inputs + selection</i>"]
    PKG --> TR["compilation trace<br/><i>every decision with its reason</i>"]
  

That is a claim about behavior. The experiment checks it.

Testing the compiler

The test plan was written down before anything ran. No model was called, and outbound connections were blocked.

The fixture has five candidates with declared sizes. These are synthetic size estimates, not tokenizer counts:

MaterialKindDeclared sizeRequiredLineage
Athe objectiveevent4yes—
Bthe source paragraphartifact6yes—
Can earlier observationevent5no—
Da claim derived from the sourceclaim3nosource-call
Eoutput from a sibling callartifact7nosibling-call

The seal forbids call lineage sibling-call. This is the isolation Part 5 depends on: an independent call must not see what its sibling produced. The budget is 16.

Arrival order does not change the selection

The same five candidates were supplied in four different orders: A B C D E, D A C B E, C B D A E, and B D A C E.

Every run selected A, B and D, for a total of 13. Every run excluded C “because budget” (13 + 5 would be 18) and E because the seal forbids call lineage sibling-call. All four produced package ID cfa87a17… and trace ID 5cb3a3ec….

Four permutations of one fixture are not a proof for all inputs. They do show that, for this input, nothing about arrival order leaked into what was selected or how it was identified.

A seal excludes sibling material, by declared lineage

At a budget of 100, so that budget excluded nothing, the sealed fixture selected A, B, C and D (total 18) and excluded E with a reason naming sibling-call.

The paired probe removed only E’s declared lineage. With nothing linking it to the sibling call, E was admitted, and the total became 25.

Both halves matter. The seal worked, and it worked on what it was told. The compiler does not look at content and infer where it came from. If lineage is not declared, the seal cannot see it.

Required material that does not fit fails; it is not trimmed

Required A and B were given sizes totalling 10 against a budget of 9, and the compiler raised ContextBudgetUnsatisfiable (“required 10 tokens exceed budget 9”) without producing a package or trace. It did not silently drop the source to make room, which would have left a model to produce an answer without ever seeing its source.

Required material that is missing fails

The request required an ID that was never offered, and the compiler raised RequiredContextMissing, naming the missing ID, without producing a package. A supplementary case required E itself while E was sealed; that also failed, with the message naming the forbidden lineage.

Recorded, then reopened

The runtime path, compile_and_record_context, compiled the fixture at a budget of 8 and appended one context.compiled event. A separate operating-system process then opened the same SQLite ledger and artifact store, and the recovered package, trace and events matched the original while the inspection added no events.

The two runtime failure cases (a required item over budget, and a missing required item) left the ledger at the same event count before and after: 5 and 5. A failed compilation appends nothing to the ledger. The request and the exception exist only because the experiment’s own harness kept them.

What the identity binds, and what it does not

A package ID looks like a fingerprint of the context. The experiment tested that impression directly, changing one thing at a time against the same fixture:

ChangePackage IDTrace ID
Actor version fixture-v1 → fixture-v2changedchanged
Arbitrary package metadata addedunchangedunchanged
A’s declared size 4 → 3 (total 13 → 12)unchangedunchanged
A’s payload rewritten under the same event IDunchangedunchanged

The last row is the important one. The objective’s text changed from “Review the paragraph using its source.” to “Changed under same ID”, and both identities stayed exactly the same.

The package ID is a selection identity. It names which IDs were chosen under which task, actor, prompt, seal and budget. It is not a content hash of the material behind those IDs.

A second probe pushed on the same seam. A required artifact ID with no bytes behind it (ghost-artifact) was offered as a candidate. It was selected, “included because required”, and the package was produced. Looking the ID up in the artifact store failed. Offering an ID is not the same as proving its bytes exist, and the compiler checks only the first.

Allowed to see is not what was sent

Now the boundary this chapter is really about.

The experiment prepared, without sending, the request that would carry the recorded package. The prepared body was:

{"model": "fixture", "messages": [{"role": "user", "content": "Review selected inputs."}]}

The selected source paragraph B, whose text begins SOURCE-B:, does not appear in it. Neither does the objective or the derived claim. The same harness then read sealed artifact E straight from the artifact store, which worked: the seal had governed the package, not the store.

Reading the adapter source confirms this is not an accident of the fixture. In that version, the OpenCode adapter builds the request text from two fields: the call’s instruction and the package’s prompt string. Nothing dereferences the selected event, artifact or claim IDs and renders their contents.

So, in CodeAI at this stage:

  • The caller determines which candidates are offered. For that inventory, seal and budget decisions are deterministic, recorded with reasons, identified by package and trace hashes, and recoverable by another process.
  • Rendered is whatever the prompt string contains. The package’s selection is not what put material into the request.
  • Received is not observable at all.

A context package is a record of selection under this context policy, not permission to access the underlying stores. It becomes a record of model input only when the rendered bytes can be bound to the prepared request, and when this experiment ran they could not.

Current practice for agents sharpens this distinction rather than dissolving it. Anthropic’s guidance on context engineering treats context as a finite resource with an attention budget, aims for the smallest set of high-signal tokens, and names techniques for long tasks: just-in-time retrieval, compaction (summarizing a conversation near the window limit and starting again from the summary), structured note-taking kept outside the window, and sub-agents that hand back condensed summaries (Rajasekaran et al., 2025).

Those techniques decide what is in an active window. The Stage 15 package and trace record something different: which offered candidates were selected or excluded and why, under which identity. They do not enumerate every durable item that might have been available. Compaction can remove detail from the active window while a durable source record preserves it elsewhere. A well-compacted transcript is still not, by itself, an account of which candidate inventory was considered or which bytes were sent.

Lampson described the general shape of this problem in 1973 (Lampson, 1973). His customer runs an untrusted service and grants it access only to the items it needs; the question is whether information can still move by some other route. His examples run from files the service may write, through the bill for the service, to the load it puts on the system. His conclusion for a trustworthy supervisor was blunt: It is necessary to enumerate them all and then to block each one.

The mapping onto context is this chapter’s, not his: the seal blocks one channel, the package, and only for lineage someone declared. The artifact store is another channel, with no block on it, while the prompt string is a third that bypasses selection entirely. Isolation claimed on the strength of the package alone is isolation of one enumerated channel.

Closing the gap

The boundary this experiment found was closed in a follow-up stage, Stage 15B, and the closure is worth reading as closely as the gap. It is opt-in: a call that sets context_render = "context-render-v1" has its selected material rendered by the runtime, and every other call behaves exactly as before.

For an opted-in call, the runtime resolves each selected ID to bytes: artifacts from the content-addressed store with their integrity checked, events as their payload in the ledger’s canonical serialization, claims as their recorded payload. An ID that cannot be resolved raises before anything is prepared, and the call records call.preparation_failed with no manifest and no attempt.

Resolved items are laid out in the same canonical order the compiler uses, each with a header carrying its content hash and length, and the instruction, context and query are composed into one input. Before any provider effect, the runtime checks that the prepared body carries exactly that text, stores the rendered bytes as an artifact, and writes the binding into the manifest. From the preserved Stage 15B manifest, trimmed:

{
  "context_package_id": "7cfdf28e…",
  "context_render_version": "context-render-v1",
  "rendered_context_sha256": "a97141e8…",
  "rendered_items": [
    {"kind": "artifact", "item_id": "eae62863…", "content_sha256": "eae62863…", "start": 178, "end": 264},
    {"kind": "claim",    "item_id": "D",         "content_sha256": "2e86fbe0…", "start": 394, "end": 681},
    {"kind": "event",    "item_id": "A",         "content_sha256": "7c73dc25…", "start": 810, "end": 859},
    {"kind": "event",    "item_id": "C",         "content_sha256": "edb5d538…", "start": 988, "end": 1029}
  ],
  "input_layout": [
    {"part": "instruction", "start": 0,    "end": 51},
    {"part": "context",     "start": 53,   "end": 1099, "sha256": "a97141e8…"},
    {"part": "query",       "start": 1101, "end": 1143}
  ],
  "request_body_sha256": "25183509…"
}

The five states now have five separate answers. The package ID is still a selection identity. Each rendered item carries a content hash, so source bytes are named. The rendered context has its own hash and stored bytes. The request body hash covers the prepared request whose text carries those bytes at the recorded offsets. And receipt by the provider is still not observable from this side.

The preregistered stage re-ran this chapter’s awkward cases against the new path. The source paragraph B’s 86 bytes appeared in the request at their recorded offsets, and sealed E was absent from both the rendered bytes and the request text. Rewriting event A’s payload under the same ID left the package ID unchanged, as before — but now the changed item was named as exactly A, and both the rendered hash and the request hash changed.

The ghost artifact failed resolution with zero transport calls and a recorded preparation failure. Permuted supply order produced identical package, rendered and request hashes, and a control call without the opt-in sent exactly the instruction and query. A separate verifier that imports neither the producer nor CodeAI passed all 59 of its declared semantic checks and rejected five targeted seeded corruptions, including B’s bytes removed and sealed E inserted with hashes and offsets recomputed to match.

The limits are part of the result. Only calls that opt in are covered, and paths that don’t — including fan-out, as Chapter 23 will show — still send the prompt string. One adapter and one API protocol were exercised. The branch that refuses a prepared body not carrying the composed text is implemented but was not exercised. Events and claims render as raw payload JSON, a provenance format rather than a prompt design. Seals are still not access control. And the transport was synthetic: the harness recorded bytes and hashes without measuring what a model does with the rendered context.

Where it is still weak

  1. Selection reaches the request only by opt-in. When the experiment ran, the adapter rendered only the instruction and prompt string. Stage 15B binds rendered selected bytes to the request for calls that opt in; every other path still sends the prompt string.
  2. The package identity is not a content hash. A payload rewritten under the same ID, a changed declared size, or added metadata leaves the package ID unchanged. For opt-in calls, per-item and rendered hashes now name the bytes beside it.
  3. Sizes are declarations. The fixture’s sizes are synthetic. On the convenience path, artifacts and claims count as zero, and event sizes are estimated from serialized length (about four characters per token), not measured by a tokenizer.
  4. Seals see only declared lineage. Remove the lineage and sibling-derived material is admitted. The compiler never infers provenance from content.
  5. Seals are not access control. Sealed material remained readable directly from the artifact store.
  6. IDs are not checked for existence. A required artifact ID with no bytes behind it was selected.
  7. At Stage 15, compilation failures leave no runtime record. Missing and over-budget compilations raise before anything is appended; the evidence of them exists only in the experiment’s harness. Chapter 16 later adds context.compilation_requested and context.compilation_failed so this state becomes durable.
  8. At Stage 15, reopen covers success only. The recorded event carries the selection and its trace, not the full offered inventory with every requirement, size and lineage. Chapter 16 takes that missing inventory as part of the working state that must survive restart.
  9. One fixture, no model. Four permutations of five candidates, and nothing here measures whether selection changes an answer.

Do this now

Forty minutes. Find out what your model actually saw.

  1. Take the last model call your system made that mattered. Write down three lists: what material was available, what was selected for the call, and what text was actually in the prepared request. If you cannot produce the third list from a record, you do not know what your client sent. Even if you can, this chapter’s final state still applies: provider receipt and processing are not observable from this side of the API.
  2. Find where your system builds the prompt. Does it render from a recorded selection, or from whatever variables were in scope?
  3. Supply the same inputs in a different order and check whether the request or any recorded identity changes.
  4. Change the content of one source without changing its ID or filename and see what notices the difference.
  5. Mark one input as required and remove it. Does the call fail, or run without it?

If you are building with an assistant:

Treat context as a compiled input, not accumulated history.
- Represent each piece of material as a candidate with kind, id, declared
  size (label it estimated unless measured), required flag and lineage ids.
- Order candidates canonically. Include required candidates first; raise if a
  required candidate is missing, sealed, or exceeds the budget. Never trim
  required material to fit.
- Exclude optional candidates by seal (declared lineage) or budget, and record
  a reason for every exclusion.
- Hash a package identity over the selection and its parameters, and a trace
  identity over every decision and reason. Say plainly that these are selection
  identities, not content hashes, unless you also hash the material's bytes.
- Record successful compilation durably and reopen it from another process.
- Then close the gap this chapter found: render the request only from the
  recorded selection, and record a hash of the rendered request next to the
  package so "selected" and "sent" can be compared.
- Test: permute inputs (identities stable), remove declared lineage (seal
  admits), required over budget and required missing (no package), payload
  changed under the same id (identity unchanged unless bytes are hashed).

Failure modes

  • Context as transcript. An order nobody chose, and content nobody selected.
  • Silent trimming. Required material dropped to fit a budget; the answer arrives anyway.
  • Order-dependent selection. The same inputs in a different order produce a different context.
  • Unrecorded exclusion. Something available was left out and nothing says why.
  • Treating a selection hash as a content hash. The source changed and the identity did not.
  • Trusting undeclared provenance. The seal only blocks lineage someone wrote down.
  • Calling a seal access control. The sealed bytes were still in the store.
  • Recording the selection, sending something else. The package said one thing; the request carried the prompt string.

What this chapter established

  • Context is an input, not a transcript. Material on the way to a model has distinct states: available ≠ eligible ≠ selected ≠ rendered ≠ received. The compiler starts from an offered candidate inventory rather than enumerating all durable material, so its exclusion trace explains what happened to offered candidates, not why an unoffered item was absent.
  • Select on purpose. Order candidates canonically. Admit required material first, and fail, never trim, when it is missing or does not fit. Exclude optional material by declared rules, with a recorded reason. Other work shows that position and irrelevant material change model behavior (Liu et al.; Shi et al.); that is the motivation, and this chapter measured no model.
  • Know what an identity names. A selection identity names which IDs were chosen under which parameters. It is not a content hash of the material behind those IDs unless the bytes are hashed too.
  • Tie the record to the request. A context package records selection under one policy path, not permission to read the underlying stores. It becomes evidence about model input only when rendered bytes are bound to the prepared request. Isolation claimed from the package alone blocks one enumerated channel, and only for lineage someone declared.
  • Some of it stays unobservable. The client can preserve the request it prepared or sent; whether the provider processed those bytes cannot be seen from this side of the API.

What CodeAI showed. In a preregistered run on one five-item fixture, CodeAI’s existing compiler behaved as specified. Four input orders gave one selection and identical package and trace identities. The seal excluded sibling-derived material by declared lineage and admitted it when the lineage was removed. A required item over budget, and a missing required item, each failed with no package. A recorded compilation was recovered by another process without adding events. Rewriting a payload under the same ID left the package identity unchanged, and a required ID with no bytes behind it was selected. The prepared request contained only the prompt string, not the selected material, and sealed material stayed readable from storage.

Stage 15B closed that gap for calls that opt in: the runtime renders the selected bytes and binds per-item hashes, the rendered hash and the layout into the manifest before any provider effect. A separate verifier passed 59 declared semantic checks and rejected the seeded corruptions described above. At that stage, calls without the opt-in still sent the prompt string; later chapters exercise and extend other paths separately.

Evidence notes

Preregistration. Hypotheses, fixture, budget and controls were written to preregistration.json at 18:52 UTC before the run; the execution record pins that file’s hash, the CodeAI version and the hash of every source file used. No model was called, and outbound connections were blocked.

Separate reconstruction. The bundle’s verifier imports neither CodeAI nor the producer. From the raw requests, results and reopened events it:

  • recomputes both identities from their fields, checks the arithmetic, the ordering and every exclusion reason
  • confirms the required failures left the ledger unchanged
  • checks persistence, and that the preregistration predates execution
  • checks every preserved file against the hash inventory

All 45 declared checks pass. This is an implementation cross-check of the recorded experiment, not evidence that the context policy is adequate for every task or that provider receipt occurred.

A verifier that always passes proves nothing, so four corrupted copies were run with the byte inventory bypassed, leaving only the semantic checks to catch them:

CorruptionClaims that failed
An exclusion reason removed, with the trace ID recomputed to matchseal-trace, H2-seal
Sealed E inserted into the selectionseal-identity, seal-trace, H2-seal
Selected event order reversedseal-identity, seal-order
The over-budget failure rewritten as a successH3

Each exited with a failure naming exactly those claims. The verifier and all four corruptions were rerun for this chapter, and both identities recomputed from raw fields for the four permutations; the results matched.

Next

The package can be recovered after a clean exit, as long as the compilation succeeded. The rest of what the decision depended on cannot. The full list of what was offered, what was required, the declared sizes, the lineage, and every attempt that failed, lives only in this experiment’s harness.

A process that picks up this work tomorrow would know what was selected, not what it was selected from, or what was tried and refused. Continuing an investigation needs more than the last answer. It needs the working state.

Continue with Restart Is Not Resume.

References

  • Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics, vol. 12 (2024), pp. 157–173. https://doi.org/10.1162/tacl_a_00638
  • Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. Large Language Models Can Be Easily Distracted by Irrelevant Context. Proceedings of the 40th International Conference on Machine Learning, PMLR 202 (2023), pp. 31210–31227. https://proceedings.mlr.press/v202/shi23a.html
  • Butler W. Lampson. A Note on the Confinement Problem. Communications of the ACM, vol. 16, no. 10 (October 1973), pp. 613–615. https://doi.org/10.1145/362375.362389
  • Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield (Anthropic Applied AI team). Effective Context Engineering for AI Agents. Anthropic Engineering, 29 September 2025. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents

Implementation and evidence sources: Stage 15 behavior is pinned by experiments/applied-ai/evidence/context-selection/2026-09-13-54e32384/ and its executed producer copy, including the preregistration, execution record, fixture, reopened ledger export, identity counterexamples, downstream boundary probe, separate verify.py, four seeded corruptions, test outputs, chapter-evidence-report.md and hashes.json. Stage 15B is pinned by experiments/applied-ai/evidence/context-rendering/2026-09-13-0f9a83b/ (preregistration, seven cases, preserved manifest and rendered bytes, separate verify.py with 59 semantic checks and five seeded corruptions, chapter-evidence-report.md). The Stage 15B mechanisms remain visible in current source: src/codeai/rendering.py (render_context, compose_model_input), CallSpec.context_render, and manifest binding plus call.preparation_failed in src/codeai/runtime.py. src/codeai/context.py contains ContextCompiler.compile_with_trace, ContextCompiler.compile_candidates, ContextCandidate, CompilationTrace, candidate_blocked_by_seal, event_lineage and _estimate_tokens; src/codeai/domain.py contains ContextPackage and Seal. Current Runtime.compile_and_record_context is a later successor: Chapter 16 changed it to record context.compilation_requested with the offered inventory before compilation and then a success or failure outcome. The historical 29 focused/full-suite counts reported for this stage are evidence-bundle results, not claims about the current suite size.