What a Compilation Proves
Read a decision trace, check what determinism buys, and separate what the compiler's construction evidence establishes from what it cannot.
The compiler has just returned a bundle of five items for the migration task. It says they were admitted for good reasons. Why should anyone believe it?
Start with what it hands back besides the bundle. This is the decision trace for one of the synthetic cases from the last chapter, the one in which a cheap reference needs an expensive resolver, at a budget of 1,400 tokens.
| Candidate | Decision | Reason | Class | Cost | Place in bundle |
|---|---|---|---|---|---|
| standing rule about generated files | admitted | mandatory | mandatory | 60 | 0 |
| the task | admitted | mandatory | mandatory | 40 | 1 |
| reference to the incident | admitted | ranked in | discretionary | 20 | 2 |
| notes on an unrelated subsystem | admitted | ranked in | discretionary | 300 | 3 |
| resolver definition | admitted | ranked in | preferred | 650 | 4 |
| full text of the incident | rejected | another form of this content admitted | discretionary | 0 | none |
Six candidates, six decisions, each with a reason. Nothing was dropped without an entry. (Places are counted from zero, and the table is shown in bundle order.)
The trace is a report of what the program did. It is not the model’s reasoning, and it is not an explanation offered after the fact by something that might be wrong about itself. Each line is what the code did at the moment it did it, written down by the code that did it. It can be read the way a build log is read, by someone who wants to know why the output looks the way it does.
What a trace says, and what it does not
Each entry names the candidate, the decision, a reason code and a short detail, the requirement class it was treated as, its relevance, the marginal cost charged to it, the budget before and after the decision, the candidate’s own dependencies, and its position in the bundle if it was admitted. A candidate that was admitted only because another one needed it also says which one.
The vocabulary of decisions is small. A candidate is admitted; or rejected by a hard gate, with the gate named; or rejected because it did not fit the budget; or because it added nothing; or because another form of the same content was admitted instead; or because a dependency of it was illegal. That is six decisions. The corpus below exercises five of them; the sixth, the dependency rejection, is written by the engine but no case in the corpus reaches it.
Entries are sorted by candidate identifier, not by the order in which decisions were made. A trace is a table to be looked up in, and the identifier is a stable key. If the compilation fails, the trace is partial: it holds the decisions reached before the refusal, not a judgement on candidates the compiler never got to.
In the trace above, the resolver’s entry says it was pulled in by the reference, and the reference’s entry lists the resolver among its dependencies. Both entries show the same budget before the decision and the same budget after it, because the reference and the resolver were admitted as one unit and the unit is what was charged: 670 tokens, not 20 and 650 separately. That is the dependency story of the last chapter, told by the compiler itself.
It has not always been able to. Two versions of the trace exist, and they should not be confused. The historical one, produced by the Python reference and preserved in its frozen outputs, has no record of what pulled a dependency in and no meaningful budget after a decision. The current one does. When the historical trace was first read closely, two of its fields turned out to carry no information. The field for the budget after a decision repeated the budget before it, and the field for a candidate’s dependencies was empty in every entry. The trace could say that the resolver had been admitted, and not why. The section on the port, below, tells how that was found and what it teaches. The version described here is the repaired one.
The trace has limits, and it is easy to over-read. It records that a decision was made, not that it was right. A trace can show that a stale observation was rejected at the freshness gate. It cannot show that the observation really was stale, because the gate was told so. The trace is honest about the compiler’s reasoning. It has no view of the world.
What it makes possible is inspection. A bundle that arrives without a trace can only be judged by reading it. A bundle with one can be audited: which candidates were considered, which lost, to what, and for what stated reason. Later, when a task fails, that is the difference between wondering whether the evidence was missing, mis-scoped or correctly excluded, and looking it up.
Same inputs, same output
The compiler is deterministic. Given the same request, the same candidates, the same policy and the same token counts, it returns the same bundle and the same trace, byte for byte.
That takes some discipline. There is no clock: the creation time comes in with the request. There is no random number, no network, no model and no reading of files. Sorting uses an explicit comparison of identifiers by character code, not the locale of whoever runs it. And the order in which candidates arrive does not matter. Handing the compiler the same pool in a different order produces an identical result, because it sorts by identifier before it decides anything. We checked that by shuffling the pool of every case in the corpus.
A bundle’s identity is a hash, and how that hash is built says what the compiler thinks a bundle is:
export function contentHash(bundle: ContextBundle): string {
const digest = createHash("sha256");
for (const item of bundle.items) {
digest.update(item.id, "utf-8");
digest.update(Buffer.from([0]));
digest.update(item.content, "utf-8");
digest.update(Buffer.from([0]));
}
return digest.digest("hex");
}
It covers each item’s identifier and content, in order. Two bundles holding the same items in a different sequence have different hashes, because they are different bundles, which is Chapter 6’s argument in ten lines. It does not cover token counts, and it does not cover the rendered text, which has a header and separators the items do not. Those get their own identities when the render is delivered, in the next chapter.
Determinism matters for a practical reason, and the reason is measurement. When an experiment produces different behaviour on different runs, there are at least three places the variation can come from: the assembly of the context, the delivery of it, and the model. If the assembly is deterministic, the first source is gone, and whatever varies later is not the compiler’s doing. It has three further uses.
Replay. A bundle from last month can be rebuilt exactly, and the trace read again, without having kept the bundle.
Differential testing. Two implementations, or two versions of one, can be run over the same inputs and compared. Any difference is a finding.
Policy comparison. Two policies can be run over the same pool, and the difference between their bundles is entirely attributable to the policy.
Failure analysis needs it too. A failing compilation that cannot be reproduced cannot be diagnosed.
Checking the compiler with something that is not the compiler
The engine assembles a bundle. A separate set of functions checks it. They take the finished bundle, the request, the candidates and the policy, and recompute the legality of the result from scratch, without trusting anything the engine recorded about itself. They confirm that the items appear in the recorded order, that the rendered cost is within budget, that no admitted candidate fails a gate or its floor, that every named source is present, that every dependency of an admitted candidate is also admitted, that no required group is half in, and that no content appears in two forms. A second function checks the result: that the trace covers every candidate considered, that it belongs to this request, and that a success does not carry a failure.
The separation earns its keep because an engine that checks itself can only agree with itself. A bug that admits an out-of-scope candidate would also make the engine believe the candidate was in scope.
The checker was not always independent. In the first version it asked the engine’s own gate function whether each admitted candidate was legal, which is asking the component under test to mark its own work. The four gate predicates are now written out again in the checker. That is duplication, and it is the useful kind: four short rules, stated twice, by code that cannot see the first statement. Tests then feed the checker bundles that break each rule in turn, and it has to reject every one.
A second function now checks what the trace claims: that an admission shows the budget it spent, and that a candidate with dependencies lists them. It would have caught the two old defects.
What the checker protects against is a mistake in the assembly: a candidate that got past a gate it should not have, a dependency that was forgotten, a budget that was miscounted. It does not protect against a mistake in what the gates mean, nor against eligibility judgements that were wrong when they were handed in.
Fourteen cases, three budgets
The compiler ships with a corpus of small synthetic cases, each built to exercise one mechanism and to tempt a weaker assembler into the mistake the mechanism prevents.
| Case | What it tempts an assembler to do |
|---|---|
| budget slack | fill a large budget with material that adds nothing |
| conflict | admit one side of a disagreement without the other |
| dependency cycle | loop, or refuse two items that need each other |
| dependency trap | count a reference at twenty tokens when it needs six hundred and fifty more |
| shared dependency | charge two references separately for a definition they share |
| mixed pool | mishandle several kinds of material under one budget |
| mandatory overflow | trim a mandatory item to fit |
| no legal representation | accept a form below the floor |
| required source unavailable | substitute a similar candidate for a missing one |
| qualification trap | admit a claim without its qualification |
| representation alternatives | admit two forms of one content, or the richest one because it fits |
| rendered overflow | trust the declared counts and overrun the real render |
| stale cheap candidate | prefer the cheaper, more relevant, out-of-date form |
| wrong scope | admit a highly relevant item from another project |
Each case is compiled at three budgets: tight, medium and roomy. Fourteen cases at three budgets is forty-two compilations. On the compiler’s own suite, thirty-four produce a bundle and eight produce an explicit refusal: two of the three budgets for mandatory overflow, and all three for each of no legal representation and required source unavailable. Every one of the forty-two matches its expected outcome. Every bundle is within budget in the units the candidates declared.
The suite has a second use, which is the last section’s subject: it is the frozen behaviour that another implementation has to reproduce.
Against the alternatives
The construction experiment asked what the compiler adds by running the same forty-two cases through six ways of building a bundle. That gives 252 compilations. The six were: dump, which fills the budget in arrival order and truncates; top-k, which ranks by relevance and takes from the top; a weighted packer, which blends relevance with penalties into one score; hard-gated greedy, which applies the four gates and then ranks; the compiler; and an oracle, which knows from a hidden ledger which items the task really needs.
Book result — compiler construction. On our synthetic construction suite (14 cases × 3 budgets × 6 strategies = 252 compilations), the deterministic staged compiler produced the correct outcome, a bundle or an explicit refusal, on all 42 cases, with no scope, freshness, authority, floor, dependency or group violations. The evidence register lists this result as Compiler construction.
The rest of the table is the contrast that gives those words weight.
| Strategy | Correct outcome, of 42 | Illegal admissions | Distractors admitted | Harmful-labelled admitted | Broke a dependency / a group |
|---|---|---|---|---|---|
| dump and truncate | 34 | 12 | 18 | 3 | 1 / 0 |
| top-k | 34 | 15 | 18 | 2 | 1 / 1 |
| weighted | 34 | 15 | 17 | 1 | 1 / 0 |
| hard-gated greedy | 34 | 0 | 9 | 2 | 1 / 1 |
| the compiler | 42 | 0 | 5 | 2 | 0 / 0 |
| oracle | 42 | 0 | 0 | 0 | 0 / 0 |
Three readings follow.
The four simpler strategies never refuse. On the eight cases where no legal bundle exists, each of them produced one anyway. On the mandatory-overflow case at the tightest budget, top-k admitted the first mandatory item, rejected the second, and filled the space it left with a small optional one, returning a 2,061-token bundle that silently lacked a constraint the task could not do without. That is the bundle that looks legal and is not.
The gates account for most of the difference in legality. Dump, top-k and the weighted packer each admit a dozen or more illegal items across the forty-two cases: material from the wrong project, stale observations, forms below their floors. The hard-gated greedy strategy admits none, because it shares the compiler’s gates. What it still gets wrong is structure. It broke one dependency and one required group, and on the stale-cheap case it recovered two-thirds of the must-have evidence where the compiler recovered all of it.
And recall is not what separates them. The weighted packer, like the compiler, recovered every must-have item on every feasible case. Judged by what it finds, a ranking looks fine. Judged by what it admits alongside, and by whether it will admit anything when it should not, it does not.
Legal is not useful
The same table carries two entries that belong beside the headline, not beneath it.
The compiler admitted five items that the evaluators’ hidden ledger marks as distractors, and two it marks as harmful. The oracle admitted none. The compiler cannot see the ledger, and legality was never a promise about usefulness. In the mixed pool, the harmful item is an instruction-shaped note with relevance 0.55 and 130 tokens. It is in scope, fresh and authorised, above the policy’s relevance threshold, and it fits. On everything the compiler can see, it deserves admission, and the compiler admits it at the medium budget and at the roomy one.
A defence of that behaviour is available: a legal admission under uncertainty is rational, and the compiler has no other information. The defence is right, and it is beside the point. The claim being made is that the bundle is legal, not that it is good. Whether a legal bundle helps is a question about a model, and the model has not come into the chapter yet.
Same behaviour, new language
The compiler was not written first in the language it is written in now. It began as part of the research code, in Python. It was then extracted into its own package, with its behaviour on all forty-two compilations captured as frozen outputs: the result, the trace, the bundle and the rendered text for each. Later it was reimplemented in TypeScript, so that it could run in the same environment as the coding agent used for the book’s live experiments.
The check that the reimplementation had not changed anything was direct. Run the new code over the same forty-two requests, candidate lists and policies, and compare everything it returns with the frozen outputs.
Book result — implementation parity. The TypeScript implementation reproduces the frozen outputs of the earlier Python implementation on all 42 compilations, with no semantic mismatches. The serialised schemas did not change across the port. The evidence register lists this result as Implementation parity, and anyone can rerun the comparison with the package’s conformance suite.
The construction experiment above ran on the Python implementation, because it predates the port. The 252-way comparison with the other strategies is Python’s output and stays Python’s output. What the port added is the forty-two compiler outcomes, reproduced.
The lesson is not that one language is better. It is about what a specification is. If a system is deterministic and its behaviour on a set of cases is frozen, then the implementation language is not part of the specification. The behaviour is. That is what let the code change language without silently changing what it decides.
Parity is not correctness
The port had to copy behaviour that nobody would have chosen. The token estimator rounds halves to the nearest even number, as Python does; JavaScript rounds them upward, so the port has a small function to imitate Python. Words are split on the same whitespace rule. Identifiers are compared by character code, as Python’s sort does.
It also copied three faults. The field that looks as though it should hold the budget after a decision held the same number as the budget before it. The dependency slot in the trace was empty in every entry. And the policy recorded how mandatory content should choose among its forms while the engine ignored the setting. All three were in the reference, and the frozen Python outputs contain the first two. Parity required the port to reproduce them, and it did.
That is the point at which the test did its job and showed its edge. It established that two implementations agree. It could not have established that they were right, because agreement is what a faithful copy of a mistake looks like. Reproducibility is not validity.
What found the faults was a different kind of check: reading the trace as an auditor would, and asking of each field what it was for. The repair was deliberately narrow. Selection was left alone: on all forty-two cases the same candidates are admitted in the same order, with the same bundle identities and the same refusals. The frozen Python outputs were left alone too, because they are a record of what the reference did and not a specification to be corrected. The trace and result schemas moved to a new version, because the meaning of fields had changed, and parity is now checked by projecting the new trace back onto the old shape. A repaired trace is compared with the historical one on everything they still share, and a separate set of tests pins what the new version adds. The policy field is now enforced, and rejects any value the engine does not support.
Parity is also weaker than it sounds for a second reason. It shows that two implementations agree on forty-two cases. It does not show that either is right on a forty-third.
What this kind of evidence is
The word proves in the chapter title is doing careful work. What has been established has a name, and it is worth having one, because the book will meet three more kinds and must keep them apart.
Structural evidence is evidence about the bundle as an artefact. The bundle fits its budget. Its dependencies are closed. No gate was violated. The trace covers every candidate. Two runs agree. It can be checked mechanically, by software, without a model and without asking anyone’s opinion.
It is strong evidence for what it covers, and the coverage has edges.
| Structural evidence shows | It does not show |
|---|---|
| under the eligibility judgements supplied, no illegal candidate was admitted | that the judgements were right |
| the bundle fits in the units the candidates declared | that it fits in a provider’s token count |
| every decision has a recorded reason | that the reasons were good ones |
| the same inputs give the same output | that the inputs describe the real repository |
| the compiler beat simpler assemblers on legality across forty-two synthetic cases | that any of them would matter on a real task |
| the two implementations agree | that both are right |
One more limit belongs in that list, because the table above hides it. The number of authority violations is zero for every strategy, including dump. No case in the corpus tempts a weaker assembler into breaking an authority rule, so the authority gate is enforced by the compiler but has never been shown to matter by contrast. The scope and freshness gates have been. The claim for authority is only that it is respected, and nothing more.
The compiler has now been given a fair hearing on what a compiler can be checked for. It is deterministic. It is validated by a second set of functions. It agrees with an earlier implementation on forty-two cases. It beats simpler ways of building a bundle on legality, and it can refuse. All of that is about the artefact the compiler produces.
None of it says that the artefact went anywhere. A bundle in the compiler’s hands is not a bundle in the model’s context. Between the two lies software that renders it, wraps it, injects it into a running session, and can fail at each step without telling the compiler. A correct compile is not a delivered one. The question the compiler cannot answer about its own output is the one the next chapter has to:
Did the bundle that was compiled actually reach the model?
References
- Project Context Compiler. Source repository, Apache-2.0 licence, first-party code by this book’s author. The engine, the independent validators, the fourteen synthetic cases with their budgets, and the frozen outputs of the earlier Python implementation used for the parity check. https://github.com/ernanhughes/project-context-compiler
- Project Context evidence register (companion repository,
evidence/README.md). Lists the identifiers, code revision and published artefacts behind each measured result. Consumed here: Compiler construction (the six-strategy comparison, the violation counts, the distractor and harmful-labelled admissions) and Implementation parity (the pinned goldens and fixtures and the recorded conformance run, including the two trace defects the goldens contain). Structural evidence only; no behavioural claim.