The Measurement Instrument
Observe the real payload sent to a model before attempting to improve it: what a read-only observer can and cannot see, how tokens are counted, and why the instrument has to be qualified before it is trusted.
Imagine a developer debugging a slow, expensive coding agent. The visible transcript shows eight user messages and eight assistant replies. It looks modest. Then they enable request logging at the harness layer and discover that the eighth model invocation carried tens of thousands of tokens: the project instructions re-sent for the eighth time, tool definitions for tools never called in this session, two file reads that overlap almost entirely, and a history that includes four failed tool calls and their stack traces in full. The user typed a few hundred words in the whole session. The model received a novella.
The numbers in that story are invented, and the story is not exotic. This gap between the imagined prompt and the rendered request is the normal condition. Every production harness assembles the invocation from many sources, and none of them is visible in the chat transcript. Developers who reason from the transcript are reasoning about a document the model never saw. The first engineering obligation is therefore not to prune, compress or retrieve anything. It is to look.
The imagined prompt
Ask a practitioner what their system sends to the model and the answer is usually the last user message plus a vague gesture toward history and instructions. The rendered request for an agentic turn routinely contains all of this:
system instructions
developer and product instructions
project instructions
tool definitions and schemas
conversation history
summaries of earlier history
files and excerpts
retrieved documents
tool results and observations
images and other media
working state (plans, task lists, notes)
current user input
Each was placed there by a concrete piece of software making a concrete decision: a harness that prepends the project file, a framework default that appends every tool schema, a history manager that concatenates instead of selecting, a compaction step that replaced six turns with a paragraph nobody inspected. Provider documentation confirms the shape. OpenAI describes the cached unit as the model’s full rendered context, including provider instructions, developer messages, tool definitions and conversation history, and adds that each request is stateless and everything is re-supplied per call. The transcript is the tip of a structure the harness rebuilds from scratch every turn.
No claim about context can be evaluated without capture. “We added the style guide to context” is ambiguous until we know where it was placed, in what form, on which turns and at what cost. “History is getting too long” is unmeasurable until history is separated from instructions, tools and retrieved material. Observation comes before every later verb in this book.
An observer that only looks
The instrument this book uses is a read-only observer attached to a real coding agent. It sits at the point where the agent has assembled the context for a model call and is about to hand it on, and it writes one record per call. It changes nothing.
Read-only is a deliberate constraint. An instrument that modifies the request while measuring it cannot establish a baseline, and baselines are the scarcest resource in context engineering. The observer never mutates the event it is looking at, and tests pin that.
Each record carries the session and the agent, the requested and observed provider and model, the kind of request, the system material, the messages, the tool definitions, the request options, the model’s limits where the host reports them, a capture time, a sequence number and timings. Two details of that record are worth having in mind.
The messages are not a flat list of strings. Each has a role and a list of parts, and the parts are of distinct kinds: text, reasoning, tool calls with their arguments, and tool results. Tool results carry no header saying which call produced them, so anything that wants to know whether two results are repeats has to compare content.
And the sequence number survives restarts. It is reserved before the record is written, so a crash leaves a detectable gap and can never hand the same number out twice. If the state is damaged and the record store cannot vouch for a session, the session is marked incomplete by rule. Order is never invented.
The observer is one mechanism of a package that also contains an injector, which Chapter 25 describes. Both are off unless explicitly enabled. An instrument that is on by default changes ordinary behaviour, and one that is off by default fails silently, so the second failure needs its own defence.
What looking refuted
The first version of the plan was written before any real capture existed. It specified a record of eleven fields, among them authority, scope and cache span, and a set of synthetic sessions to test against. When the observer was finally run against a real session, in a throwaway repository, several assumptions failed.
The messages were not shaped the way every synthetic fixture had assumed. Every fixture had the wrong form, so a real session would have been silently misread, its repeated tool outputs never matching. The fixtures were rebuilt to the real shape, and a record in any other shape is now refused as incomplete.
Tool-call arguments turned out to be captured, which meant edited files, test runs and repeated calls could be read from what the agent actually did rather than from proxies.
The call meant to read the model’s limits did not exist and failed silently, so every limit had been recorded as unknown. It was replaced with the host’s model registry, and a model configured without a limit is now recorded as having none, not as zero.
Provider-reported usage, cost, latency and cache reads turned out not to be invisible: the host keeps them in its own database. They can be read afterwards, read-only, and joined to a captured request only when the count of assistant messages equals the count of captured requests. Values obtained that way are labelled as joined, not observed.
And the record does not carry authority or scope at all. The observer can see where an item sits and what it says. It cannot see who placed it or what it claims to govern, because the host does not say. Those fields, which Chapters 3, 19 and 21 rely on, are not something capture can supply. They have to come from whatever builds the candidates later.
The lesson generalises. An instrument specified without contact with the thing it measures will be wrong in ways nobody has thought of, and the way to find out is to run it once, early, on something small, and to record every place the specification was refuted.
How tokens are counted
A token count is not a fact about a request. It is a measurement, and the same request has at least three.
A count can be declared: the builder of a candidate says how many tokens it has. It can be estimated by a rule, such as words times a constant. Or it can be reported by the provider after the fact.
These disagree, and by more than anyone expects. In the observer’s calibration, on one session with one model, a word-based estimate ran roughly a third below what the provider reported for the same requests, while bytes divided by four ran much closer. That is one session and one model, and it does not say how either behaves in general. It does say that a word estimate should not be the number a budget rests on without knowing which count it is.
Two consequences carry forward. Every count recorded should say which of the three it is. And when the compiler of Chapter 23 reports that a bundle fits its budget, it means in the units the candidates declared, which is a statement of a different kind from one about the provider’s tokenizer.
The tokenizer itself changes across models, so cross-model comparisons of token counts are approximate. Anything that cannot be observed is recorded as unavailable, never as zero.
Baselines before interventions
With the instrument defined, the chapter can say what it means to have measured a session. A baseline is a set of numbers computed over captured records before any intervention, and the book asks for eight.
- Total rendered input per invocation, across the session. The single number everything else decomposes.
- Tokens by category and source. Instructions versus history versus tools versus files is the minimum useful cut.
- Repeated material. Tokens substantively identical to tokens already present earlier in the same invocation or re-sent unchanged across turns.
- Growth over turns. Total and per-category counts against turn number. Agentic sessions grow; the shape distinguishes healthy accumulation from a leak.
- Stable prefix size. The leading span identical to the previous invocation, where observable. This predicts cache behaviour and exposes churn near the front.
- Tool-definition overhead. Tokens consumed by tool names, descriptions and schemas before any tool is called.
- Tool-output growth. Tokens from tool results and their share of the bundle over time. Tool output is the fastest-growing category in coding sessions.
- File and retrieval contribution. Tokens from project files, excerpts and retrieved documents, separated from history, so that “the model knew the codebase” can be checked against what was shown.
Repetition, stability and age are not stored in the record. They are computed afterwards by comparing consecutive records byte for byte.
Each later intervention in the book (deduplication, pruning, compaction, externalisation) must move at least one of these numbers and then show, by a controlled comparison, that behaviour survived or improved. A chapter that claims savings without a before and after on these baselines has made no claim.
A worked capture
Here is turn five of a debugging session, in the form a capture would take. The categories and token counts are illustrative and are not a measurement; a real table would carry measured counts, labelled with which kind of count they are.
project instructions (stable, turns 1-5) 1,900
tool definitions (stable, 14 tools) 6,200
conversation history (turns 1-4, growing) 9,800
prior summary (written at turn 4) 900
re-read source files (2 files, 80% overlap) 7,100
tool results (test output, search hits) 4,900
current user message 120
environment state 480
Three things jump out before any intervention is contemplated. The user’s message is under half a per cent of the bundle. The two largest categories, history and files, are also the fastest growing, so projecting the growth predicts the bundle doubling within a few turns. And nearly a third of the file tokens repeat material already present in the same invocation, while the tool definitions have been re-sent unchanged five times. None of that was visible in the transcript. All of it is now a row in a table, which means it can be tracked, budgeted and tested.
Stable and dynamic
Of the derived properties, stability deserves emphasis because it connects observation directly to economics. A stable item is byte-identical across turns: the project file nobody edited, the tool schema nobody changed. A dynamic item is regenerated, appended or replaced: new tool results, fresh retrievals, the latest user message, a re-rendered timestamp.
Stability matters for two independent reasons. The behavioural reason is that stable material is the background of every turn, always present, rarely examined, and able to steer silently. A stale project rule exerts influence precisely because nobody re-reads it. The economic reason is caching: providers reuse computation over matching prefixes, and stable leading spans are what make that possible. OpenAI’s documentation states the prefix-match requirement explicitly and warns that summarising or truncating resets reuse from the first changed token. A bundle whose early spans churn every turn, perhaps because a timestamp is rendered into the opening lines, pays full price for stability it could have had for free.
Repetition is its companion. A bundle can grow without repeating, by accumulating new observations, or repeat without growing, by re-sending the same preamble. The two call for different responses, and the instrument keeps them separate.
Measuring is not improving
measuring context ≠ improving context
Every technique the book will test is, at this stage, an unexamined proposal. Showing that history can be summarised is not showing that it should be. Counting duplicate tokens is not removing them.
There is a temptation to skip this stage on the grounds that the waste is obvious and the fix is obvious. Sometimes both are. The instrument still matters, for two reasons. Obvious waste is often load-bearing: the duplicated instructions may be the only reason the model obeys them, and removal must be tested, not assumed. And the instrument is what makes the test possible. Without per-item records an intervention’s effect is argued from anecdote; with them it is argued from numbers.
Numbers need interpretation, and the most informative reading of a baseline is the shape of growth. Steady accumulation, with each turn adding a roughly constant amount, is the healthy case. Accelerating growth, where tool outputs trigger further tool calls whose outputs trigger more, marks a loop that may be productive investigation or pathological thrashing. A step change, where the total jumps, usually marks a large file read, a paste or a compaction, and each deserves to be named by category and not absorbed into an average. The diagnostic question is always the same: which categories drive the shape, and would the behaviour survive their absence? A baseline does not answer the second half. It tells us where to run the ablation.
What the baselines are expected to show
The book does not yet have a corpus of real sessions to answer that. What follows are predictions, stated as hypotheses to be confirmed or refuted by the frozen measurements, and none of them has been tested.
Book hypothesis. In real multi-turn coding traces, repeated and stable-but-unexamined material will make up a large share of rendered tokens, and the user’s own words will be a small minority.
More specifically: tool definitions are a large fixed cost from the first turn; history and tool output dominate growth; project instructions are re-sent unchanged where re-sending buys nothing; and overlapping file reads recur because each retrieval is issued without consulting what the bundle already holds. If real bundles turn out lean, non-repeating and user-dominated, several later chapters lose their motivation, and the report will say so. A baseline that cannot surprise us is not a baseline.
Book hypothesis. Small amounts of volatile material placed early in the rendered order will account for a disproportionate share of cache-prefix invalidation.
This follows from the prefix-match mechanics in the provider documentation, not from any new claim, but it needs confirmation at the level of sessions: how often do timestamps, session identifiers or reordered blocks break an otherwise stable prefix?
The instrument has been run against a deliberate calibration session in a throwaway repository, to test the tool. That session is excluded from the corpus by construction and feeds no claim. The corpus of genuine sessions is empty. Chapters that make predictions about real traces say so where they make them, and this is where the book first states the gap in full: it holds controlled evidence on synthetic tasks, and no evidence yet about how often any of these problems occur in real work. The book calls evidence of that last kind ecological. Chapter 26 sets it beside the three other kinds this book collects, and Chapter 27 says how much of each exists.
What the instrument cannot see
The observer sees the assembled context at one place: the host’s model-context hook, before the runtime translates what it has into whatever a provider expects. Everything after that point is invisible to it.
model-context observation ≠ provider-wire capture
Still unobserved: the state of the repository, the provider’s cache decisions, anything the provider adds, the order in which the provider renders the request, whether reasoning material reaches the provider, how the system block is split, and the relations between a parent agent and its subagents. Some of these are recoverable with other instruments, and provider-reported usage, joined afterwards, shows how far one can go. Others are closed.
The set of things the instrument does not see has to be a first-class part of every result that uses it. A capture that silently drops what it cannot parse is worse than none, because its totals reconcile and its blindness is invisible. Server-side state, such as provider-stored conversation items or compacted representations, is marked opaque and not invented. Encrypted or redacted spans, such as reasoning items replayed opaquely, are recorded as present but unreadable. Captured context can contain sensitive material, so it stays local and is never committed, and the corpus needs the same handling as any production log.
Capture also changes incentives. Once a team watches token counts, harnesses get tuned to the dashboard: categories shrink cosmetically, material moves to unlogged channels, totals fall while outcomes do not improve. The defence is the book’s standing rule that no optimisation counts without a controlled behavioural comparison. The instrument reports. The experiment decides.
The instrument has to be qualified
One more property has to be established before any of this can be relied on, and the book returns to it later at length. An observer that is not running looks exactly like an observer that saw nothing.
A record store that is empty proves nothing about model activity. It most likely means the observer was never switched on for the process that ran the model, and the failure is silent. So the observer is not trusted because it is installed, or because its version is right, or because its files match. It is trusted after it has been shown, in the run that matters, to capture a known event. Chapter 25 tells what happens to an experiment that skipped this step.
The question this creates
Once the rendered request is visible and counted, one fact dominates the tables and demands explanation. In session after session the largest categories will be material the user never typed: instructions written by the vendor and the product team, tool schemas generated from code, project files injected by the harness, histories and summaries accumulated without any explicit act. The user’s words are a minority input to the computation that answers them.
If we actually inspect a production request, how much of what the model receives did the user never explicitly type?
References
- OpenAI. “Prompt caching.” Official documentation, verified September 2026. Full rendered context as cache unit; prefix-match requirement; stability economics. https://developers.openai.com/api/docs/guides/prompt-caching
- OpenAI. “Conversation state.” Official documentation, verified September 2026. Stateless requests; re-supply of history per call. https://developers.openai.com/api/docs/guides/conversation-state
- Anthropic. “Writing effective tools for agents — with agents.” Published 11 September 2025. Tool definitions and responses as context costs; token-efficient tool design. https://www.anthropic.com/engineering/writing-tools-for-agents
- Liu, N. F., et al. “Lost in the Middle: How Language Models Use Long Contexts.” TACL 2023; arXiv:2307.03172. Peer-reviewed. Motivation for recording position. https://arxiv.org/abs/2307.03172
- Project Context OpenCode. Source repository, Apache-2.0 licence, first-party code by this book’s author. The read-only observer, its record format and its calibration notes. https://github.com/ernanhughes/project-context-opencode