The Measurement Instrument
Define, validate, and build the instrument we will use to measure whether an AI system actually remembers.
Chapter 1 ended with a test. Does the past change what the system does now, and does it change it for the better?
That is easy to state and hard to run. This chapter builds the thing that runs it.
Two words will recur for the rest of the book, and they must not blur together. The memory system is the thing that remembers. The instrument is the apparatus that decides whether it does. They are separate products, and this book builds both.
Neither is trustworthy alone. A memory system with no instrument is a demo. An instrument with no memory system has nothing to indict.
The instrument is not finished here. This chapter establishes version 0.1: the measurement model, the task families, the corpora, the baselines and controls, the failure taxonomy, and the first runnable scorers. Later chapters extend it as new capabilities arrive. Decisions earn decision tests. Provenance earns evidence-chain scoring. Time earns supersession tests.
The two products grow together. The memory system becomes more capable, and the instrument becomes better at exposing where it is not.
The question
The chapter revolves around a single question: how do we know whether an AI memory is any good? That question separates into four, and the instrument must answer all of them.
What do we measure? Which observable properties actually make up memory quality. The list is longer than most benchmarks assume, and a later section works through it.
How do we measure them? Which tasks, fixtures, controls, and metrics expose those properties without also rewarding fluency, lexical overlap, or benchmark exploitation?
How do we score them? What counts as success, partial success, failure, abstention, hallucination, regression, or harmful memory — and at what granularity?
How do we know the measurement itself is valid? This final question is the one most benchmarks skip. An instrument whose scores rise while downstream behaviour stagnates is measuring something, but not memory. Validity must itself be tested, and the chapter returns to it near the end.
The instrument is a pipe with one moving dial: the memory condition. Within a matched comparison, everything else stays fixed so the dial’s effect is observable:
flowchart TD
HIDDEN[Hidden canonical state] --> REND[Rendered project history]
REND --> COND[Memory condition]
COND --> RUNNER[Fixed reader or agent]
RUNNER --> OBS[Observable answer or action]
OBS --> EVAL[Evaluation]
The diagram makes the discipline visible. The evaluator sees the hidden canonical state. The system under test sees only the rendered history, plus whatever memory the condition hands it. Within a matched run, the reader stays fixed so that a model change cannot masquerade as a memory effect. Later reader-transfer experiments vary that dimension deliberately.
Why demos are not evidence
A typical memory demo proceeds like this: the builder loads an interesting history, asks a question whose answer is known, and shows the system’s fluent response. The audience nods. Nothing about this procedure distinguishes remembering from paraphrase.
Four specific defects make demos unreliable as evidence.
First, the builder selects the case. Cases where the system happens to work are shown; cases where it confuses a proposal with a decision are not. Selection bias is invisible to the audience.
Second, the query is asked once, with a convenient phrasing. Whether the system survives paraphrase, temporal reframing (“what did we use in April?” versus “what should new code target?”), or a query that resembles a rejected idea is never tested.
Third, the answer is judged by fluency. A response that names the right topic, cites a real session, and sounds confident passes inspection even when it reports a superseded decision as current or attributes the decision to the wrong reason.
Fourth, nothing is rerun. When the next mechanism is added, there is no fixed set of old cases to check for regressions. Improvements are claimed; breakage is discovered by users.
The instrument exists to constrain all four degrees of freedom at once: the corpus and queries are fixed before the system is built, the labels are hidden from the system, the metrics are defined per question, and every mechanism reruns the same cases.
What exactly are we trying to measure?
Chapter 1 ended with an evidence ladder and the warning that a system can pass an early rung while failing a later one. Chapter 2 converts that warning into a measurement model. The v0.1 dimensions below are a starting taxonomy, not a final one; later chapters may split or merge them as failures demand.
Preservation. Did the relevant historical information survive? This is the lowest rung and the easiest to verify mechanically.
Retrieval. Can the correct evidence be located? Standard ranking quantities apply here as established background: recall at k, precision at k, MRR, and nDCG where appropriate.
Reconstruction. Can the system infer what actually happened — proposal versus decision, accepted versus rejected option, conclusion versus discussion? Retrieval success does not imply reconstruction success, and the book predicts the two will diverge.
Book hypothesis. Decision reconstruction will fail on a substantial class of cases even when retrieval of the relevant discussion remains strong, because the retrieved passages do not distinguish proposals from outcomes.
Provenance. Can the system identify why the conclusion exists? Measurements include evidence recall, evidence precision, support-chain correctness, and the rate of unsupported rationales. A plausible rationale with incorrect provenance scores as a failure.
Temporal correctness. Can the system distinguish what was true then from what is true now? Measurements include current-state accuracy, historical-state accuracy, supersession correctness, and validity-interval correctness.
Epistemic correctness. Can the system preserve how something is known — evidence, observation, inference, hypothesis, speculation, unresolved conflict, rejected idea, superseded belief? The characteristic error is promotion: a hypothesis reported as fact, a superseded belief reported as current.
Groundedness. Does the output introduce historical claims unsupported by the available record — fabricated sources, incorrect provenance, confident recall of events that never occurred?
Abstention. When the history does not determine an answer, does the system say so? Refusing to manufacture an answer is part of memory quality.
Context selection. When memory must fit a bounded context, were the useful memories selected and the harmful ones excluded? Measurements separate omission of critical memories from admission of stale, irrelevant, or contradictory ones.
Behavioural influence and utility. Did the retained past change what the system did, and did the change improve the outcome? These are the Chapter 1 counterfactual made runnable, and the hardest rung to instrument.
Harmful memory. Did memory make the system worse — through stale, irrelevant, contradictory, or misleading recall, excessive context, or false confidence induced by remembered material? Harm deserves first-class measurement, not an appendix.
Cost. Tokens, latency, storage, model calls, context usage. Memory quality is not independent of resource usage, and a system that remembers well at unbounded cost has shown less than it appears to.
No single number spans these dimensions. A system can show excellent retrieval with poor temporal memory, or excellent provenance with terrible behavioural utility, and an aggregate would hide both facts. The instrument reports a scorecard; any aggregate stays secondary and optional.
A task is not a metric
Much evaluation confusion comes from collapsing five distinct things into “a memory question.” The instrument keeps them separate:
- Task. What event store should the new service use?
- Ground truth. The hidden ledger establishes PostgreSQL as the current target, decided in
adr-007, superseding SQLite. - Evidence. The ledger’s supporting entries: the contention experiment, the incident report, the decision record.
- Expected behaviour. Use PostgreSQL and optionally explain the relevant reason.
- Measurements. Correct current decision, correct evidence, no superseded recommendation, no fabricated rationale, appropriate use of memory.
- Scores. Individual metric outputs, each traceable to the above.
This separation is enforced in code. A measurement task carries the prompt and a reference to the history the system may see, while every expected_* field is evaluator ground truth the system must never receive:
task = MemoryTask(
task_id="decision-event-store",
family="decision",
prompt="What should new event-store services use?",
history_ref="controlled-v0.1",
expected_sources=("adr-007",),
expected_state="PostgreSQL",
superseded_options=("SQLite",),
)
A task family fixes the rest: input shape, expected evidence, allowed uncertainty, expected behaviour, metrics, and failure labels.
The book’s six questions supply the first six families — locate, decide, justify, track truth through time, track unfinished work, and select what matters now. These are measurement families, not definitions of memory. Later chapters add more as new capabilities arrive.
What prior work measures
The papers already cited by this book were inspected for one question each: how did they know their system remembered better? The full research matrix lives outside the published chapters; what follows is what the instrument borrows and what it refuses.
LongMemEval separates information extraction, multi-session reasoning, temporal reasoning, knowledge updating, and abstention, with evidence-position labels in long distractor histories. The instrument borrows the ability split and the abstention task, and refuses judge-only scoring: every LongMemEval correctness verdict comes from a single proprietary model judge, which makes the judge a single point of validity failure.
LoCoMo evaluates very long conversational histories with adversarial items the system should resist, and compares retrieval over raw dialogues, observations, and summaries. The instrument borrows adversarial abstention items and the observation-level retrieval comparison, and refuses lexical overlap as a sufficient measure: word-match scores cannot separate recall from understanding.
LongBench contributes two controls the instrument adopts directly. A no-context condition separates what the model already knew from what the history supplied, and compression-versus-capability comparisons test whether a smaller representation preserves what matters. LongBench is otherwise refused as a memory instrument: it measures long-context use, not updating, abstention, or behaviour change from history.
AgentBench evaluates agents through interaction rather than prose, with a finish-reason taxonomy that distinguishes genuine completion from context-limit, invalid-action, and task-limit endings. The instrument borrows interactive evaluation and finish reasons for its later behavioural families, and refuses cohort-dependent aggregates and truncated histories that handicap long-horizon behaviour.
The memory systems of Chapter 1 supply a different lesson.
MemGPT and MemoryBank also motivate caution around small samples and circular evaluation. Circular here means testing on histories produced by the same model family that answers the questions.
Generative Agents shows the value of ablations that remove memory, reflection, and planning separately, so each component has to earn its keep.
The instrument takes the ablation discipline and the suspicion of self-graded histories from these systems. It also makes false memory explicitly scoreable rather than rewarding only successful recall.
Prior work reports (not book results): benchmark-scale memory evaluation currently leans on model judges, lexical overlap, and short histories. The instrument is designed so that each of those shortcuts has a control that exposes it.
Known truth and real history
The benchmark ingests two kinds of history because no single corpus answers both questions the book needs answered: does the mechanism work when ground truth is known exactly, and does it still work when history was not constructed for the test?
The controlled corpus: mechanical ground truth
A generator produces synthetic project histories from a hidden ledger. The ledger is authoritative. It records which events occurred, which proposals became decisions, which facts held during which intervals, what superseded what, and which tasks were left open.
A rendering step then turns the ledger into ordinary project debris: session transcripts, commits, document revisions, issues, decision records, experiment notes.
The memory system sees only the debris. The evaluator sees the debris plus the ledger, and derives the expected answers mechanically from it.
That separation is the core control of the whole benchmark. If the system could read the ledger, the measurement would be circular — rewarding systems that agree with the labeller rather than systems that reconstruct what happened.
One piece of terminology, to keep the confusion from ever starting. The hidden representation is the Evaluator Ground-Truth Ledger. It is evaluation ground truth, not system architecture. Nothing about its schema tells a memory system how to represent anything.
A minimal ledger entry, shown here as an illustration of the evaluation representation, carries the fields the spec requires. The canonical decision entry, simplified for illustration, is evt-205:
event_id: evt-205
kind: decision
topic: event-store target
content: New event-store work should target PostgreSQL.
actors: [m.okafor, j.lindqvist]
decided_at: 2024-07-11
valid_from: 2024-07-11
valid_until: null
supersedes: [evt-201] # the earlier SQLite decision
supported_by: [evt-202, evt-203, evt-204] # contention, benchmark, incident
source: [adr-007]
status: current
The ledger keeps production state separate from decisions, and the two carry different dates. The decision to move to PostgreSQL lands first. The actual cutover happens eleven days later.
For those eleven days, the decision says PostgreSQL and production still says SQLite. Both are true.
So two questions that sound almost identical resolve against different records. What should new code target? reads the decision. What runs in production? reads the production state. A system that collapses both into one undifferentiated state cannot represent both answers correctly at the same point in time.
The spec enumerates a long list of ledger fields — claim, proposal, decision, rationale, validity interval, supersession, task, source. That list does not dictate the memory system’s internal schema. It exists so the generator can plant the situations the book needs:
- A discussion that ends in a decision.
- A proposal that ends in a rejection.
- A decision that is later reversed.
- A fact that a later fact supersedes.
- Two similar conversations with different outcomes.
- A summary that repeats a truth which has since expired.
The real corpus: ecological validation
The second corpus is a genuine long-running project history, used with a smaller manually adjudicated query set. Every answer retains evidence references, and where adjudicators disagree, the disagreement is preserved rather than resolved by fiat.
The real corpus cannot supply mechanical ground truth at the same scale. Adjudication is expensive, and real history rarely carries clean labels for validity or supersession.
What it supplies instead is resistance to generator artefacts. Any synthetic generator bakes in assumptions — how decisions get phrased, how far apart discussion and decision sit, how explicit the timestamps are. A mechanism that quietly exploits those regularities will look strong on the controlled corpus and collapse on real history. The real corpus exists to catch exactly that.
The two scores are never merged into one number. A mechanism earns its place by improving the controlled measurement for its intended failure class without unacceptable regressions elsewhere, and by surviving contact with the real corpus. Either result alone is insufficient.
Baselines and controls
Every experiment includes the applicable subset of the spec’s baseline ladder: no memory at all, lexical retrieval, embedding top-k, embedding plus reranking, the best system so far, the candidate system with the new mechanism, and an ablation with the mechanism removed or neutralized. The ablation matters more than it may seem. Without it, an improvement can be attributed to the mechanism when it actually came from a changed prompt, a larger context budget, or a different chunking policy that rode along with the patch.
The ladder needs adversaries as well as baselines. Each one answers a specific suspicion:
- Random history, matched for size. Does any history help, or does the right history help?
- Scrambled history, same token count and formatting, content destroyed. Was the improvement caused by content, or just by a longer context?
- Oracle memory, supplying exactly what the task needs. An approximate ceiling for selection, independent of retrieval.
- Stale and distractor memory, injecting plausible-but-superseded or relevant-looking-but-wrong material. Does the system prefer current truth over familiar text?
One more control matters where a claim is meant to hold across model capability. Reader transfer freezes the mechanism, task, context, and grader, and varies only the reader.
The logic is worth stating plainly. If a memory effect survives a much stronger reader, that is evidence that the retained history contributes information the stronger reader still uses. If the effect vanishes, the earlier gain may instead have depended on an interaction with the weaker reader’s capabilities. Reader transfer distinguishes those possibilities without assigning either outcome a single cause.
The behavioural claim itself is comparative. Conceptually, every memory verdict has the shape:
without_memory = run(task_set, memory=None) # schematic
with_memory = run(task_set, memory=memory_system) # schematic
delta = score(with_memory) - score(without_memory)
Those three lines hide a great deal. A real frozen run pins everything the result could depend on: corpus version and generator seed, task-set version, model identity, prompts, embedding and reranking models, chunking policy, retrieval and context budgets, memory configuration, code commit, and grader version.
Change any one of them and it is a new run, not the same one. That is what turns “rerun the same cases” from a slogan into a reproducible protocol. The manifest records the dependencies needed to reconstruct the reported run.
Failure attribution uses the spec’s categories: ingestion, encoding, storage, retrieval, ranking, temporal reasoning, context assembly, downstream reasoning, and evaluator defect.
That last category is easy to overlook and genuinely important. Sometimes the harness is wrong — the generator emitted an ambiguous timestamp, or a query admits two answers the ledger would both accept. That fix belongs to the benchmark, versioned as a new release. It must never be quietly folded into a claimed system improvement.
What counts as failure
A failed answer never produces a bare failure verdict. Each failure is attributed to the layer it indicts, because a measurement should tell you which layer to repair.
The v0.1 taxonomy keeps these pairs apart, among others:
- Never stored, versus stored but never retrieved.
- Wrong evidence retrieved, versus right evidence read wrongly.
- A proposal reported as a decision, versus a rejected option reported as accepted.
- Stale state, versus a supersession that was missed.
- Missing provenance, versus false provenance.
- An unsupported claim, versus a failure to abstain.
- Context omitted, versus context that distracted.
- Correct memory retrieved but unused, versus memory that actively made things worse.
Two more labels cover the cases where the harness is at fault rather than the system: evaluator defect, and ambiguous ground truth.
Two scorers show the granularity. Provenance scoring separates what was found from what was relevant: source recall (fraction of the ledger’s required evidence retrieved), source precision (fraction of the retrieved material that was relevant), and an unsupported-source rate (fraction of cited sources absent from history).
Temporal scoring keeps the two directions of truth apart. A historically accurate answer to a current-state question is stale, not partially correct: current-mode scoring matches expected current state or records stale state, and historical-mode scoring matches expected historical state or fails.
Two examples show why the label matters more than the number.
A system that answers “SQLite” when asked what new code should target receives MISSED_SUPERSESSION — not a low similarity score. A system that answers “PostgreSQL has always been the choice” when asked about the earlier period receives a historical-state failure, even though the sentence names the right technology.
The failure label localises the layer under diagnosis. Later chapters extend this scoring without changing its meaning, adding temporal-role accuracy for trajectories and status accuracy for open loops. Each is reported on its own; no aggregate score is defined.
Evidence and hallucination
Take a memory system that says: we decided PostgreSQL because of benchmark X. That one sentence has to be scoreable at several levels:
- Did PostgreSQL actually become the decision?
- Was benchmark X real?
- Was it actually evidence for that decision?
- Was necessary evidence left out?
- Was a rejected argument presented as the rationale?
- Is the claim still current?
This is where the author’s companion volume on hallucination meets this book, and the instrument inherits its evidence discipline.
The Hallucination problem asks whether a statement is supported. The Memory problem adds whether this was the correct historical or current state to bring forward into the present. The instrument keeps both questions and never lets one answer stand in for the other.
Concretely reused machinery, adapted from retrieved documents to retained history:
- Claim-edge granularity. Long responses score per claim-evidence edge, not per response. One supported claim does not launder an adjacent fabrication.
- Three evidence sets kept distinct. History available in principle, history retrieved into context, and history cited for a specific claim are three different sets. Conflating them hides the most common failure: the right past was present and the system still argued from the wrong part of it.
- Sensors are not verdicts. Scorer outputs stay typed records — entailment-style containment, trace status, source freshness — rather than averaged scalars. Policy decides what the numbers mean; measurement only reports them.
- Answerability and abstention. Each task declares whether the history determines an answer. Counterfactual pairs probe the boundary: add the decisive history and the system should answer; remove it and the system should retrieve further or abstain.
- Adversarial grading. Paraphrase must not move a score; negation, removal, role swaps, and temporal reversals must. Difficulty ladders run from random mismatch to structural inversion.
- Contamination accounting. Taint-escape rates, exposure before containment, and descendant counts per root apply to benchmark material as much as to stored claims: a leaked query set is a contaminated instrument.
- Every number with its provenance. The scorecard records where each figure was measured, with what grader, and what it does not establish — the Evidence Ledger pattern applied to the instrument itself.
Does the score mean anything?
An instrument that cannot be falsified is a ritual, not a measurement.
At this point in the investigation, the book therefore pre-registers a study that could indict its own benchmark. It is called the metric-behaviour bridge. Across several memory systems, per-question scores will be compared with downstream behavioural improvement on matched tasks.
Here is the outcome that would hurt. If decision accuracy, current-state accuracy, and provenance precision all rise while downstream behaviour stays flat, the verdict falls on the metrics, not the systems. The instrument would then need revision. That outcome is explicitly allowed.
Three further requirements govern every behavioural experiment in the book:
- A positive control. An oracle-memory condition must separate from the naive baseline. Without that, a null result confounds tasks that cannot tell the difference with systems that do not differ.
- Split discipline. Thresholds tune on development splits and evaluate on held-out tasks, with the split recorded in the manifest.
- A fluent-summariser baseline. Retrieval plus paraphrase, with no intended behavioural use. Metrics claiming to measure memory rather than text should distinguish it from conditions where retained history changes the task outcome. If fluent paraphrase consistently earns the same score, the metric is measuring too little.
The v0.1 implementation
At this stage of the investigation, the runnable instrument consists of task, observation, and manifest representations, plus deterministic scorers for source recall and precision, decision exactness, current-state and historical-state accuracy, supersession, abstention, and unsupported-source detection. Each emits a typed observation with a failure attribution attached.
Ranking metrics, support-chain validity, calibration, epistemic promotion errors, behavioural deltas, harm rates, and cost accounting are still interfaces with empty scorecard fields. Later chapters either implement them when a mechanism requires them or leave the gap explicit.
Model-judged scoring is represented but is not part of v0.1. Where later experiments require judgement, the grader’s identity, version, and configuration live in the manifest — never blurred into a mechanical score.
The implementation, including a runnable six-task demonstration against three canned systems, is described in the repository’s benchmark notes. Nothing in this section is a book result.
Controls: what the harness must prevent
The chapter’s least glamorous section is its most consequential. A benchmark that leaks answers measures the leak, not the memory. The spec therefore requires the following controls, which the harness must enforce as the corresponding fixtures become runnable.
Answer leakage. The expected answer, or a paraphrase of it, must not appear in the rendered artifacts except where the ledger intends it. In particular, decision records that restate the ledger verbatim would let a system score on decision questions by pure extraction. The generator must render decisions the way projects do: sometimes crisp, sometimes buried in a session, sometimes split across artifacts.
Timestamp leakage. If every decision carries an explicit ISO timestamp in a fixed template position while discussion never does, a system can learn “prefer timestamped passages” instead of learning authority. Timestamps must be distributed the way real projects distribute them: present in commits, inconsistent in sessions, occasionally absent or relative (“last Thursday”).
Template leakage. If each scenario renders from one template with fixed phrasing, systems learn the template. The generator needs multiple surface realizations per scenario type and, where feasible, paraphrase passes that preserve ledger semantics while varying wording.
Lexical shortcuts. If the query shares distinctive vocabulary with exactly one artifact, retrieval succeeds without understanding. The initial scenarios deliberately include similar discussions with different outcomes and shared vocabulary, so that lexical overlap is necessary but not sufficient.
Query duplication. Queries must not repeat ledger labels verbatim in ways that let the system match strings rather than reconstruct state. Query phrasing is drawn from a separate pool from ledger content.
Evaluator privilege. The evaluator uses the ledger; the system must never see it, including indirectly through file paths, generation metadata, or deterministic ordering that correlates with scenario type. Frozen fixtures should be audited for such channels before any run counts.
Memorised patterns. Because the harness will eventually run against models pretrained on public text, frozen query sets must be treated as contaminated once published: later runs report the publication date of the query set relative to model training cutoffs, and unreleased held-out variants are kept for confirmation runs.
At this stage of the investigation
Three things described above are not yet runnable here: the controlled-world generator, the adjudicated real corpus, and the metric-behaviour bridge study. The normative contract is the versioned benchmark spec in the repository. Later chapters build the controlled behavioural fixtures and the real-project transfer set, then run the bridge in Chapter 12; that later completion does not turn this early demonstration into a book result.
Before any chapter can report a book result, the harness must produce a versioned controlled corpus with hidden-ledger labelling, a versioned query set, frozen run manifests for the baseline ladder, and a failure-attribution procedure.
Two outcomes are worth naming in advance. Strong locate recall paired with weak decision accuracy would support the book’s central progression. Decision accuracy tracking locate recall closely would instead suggest that retrieval quality dominates, and that the representational machinery of later chapters needs a stronger justification than this book has yet given it.
A living instrument
The v0.1 instrument is deliberately incomplete. Each later part of the book earns new instrumentation alongside new memory machinery:
Chapter introduces decisions
→ instrument gains decision-reconstruction tests
Chapter introduces provenance
→ instrument gains evidence-chain scoring
Chapter introduces temporal memory
→ instrument gains supersession and current-state tests
Chapter introduces open loops
→ instrument gains task-state metrics
Chapter introduces context assembly
→ instrument gains relevance and distraction measurement
Chapter introduces consolidation
→ instrument gains transfer and compression measurements
Chapter introduces behavioural intervention
→ instrument measures downstream behavioural improvement
Capstone
→ instrument verifies end-to-end trace continuity and integration
If a planned component depends on machinery not yet introduced, the instrument carries the interface and an experiment slot, and the scorecard marks the field pending. That discipline is visible in the demo output already: pending fields are explicit obligations that a later chapter must either fund with a mechanism or remove honestly.
What remains unsolved
The contract is written and the first scorers run. At this point, three obligations remain open: the controlled-world generator, the adjudicated real corpus, and the metric-behaviour bridge that could indict the instrument itself. The reader will see those obligations discharged progressively, culminating in Chapter 12’s matched behavioural interventions and real-project transfer.
Others are open too — adjudication protocols, held-out rotation, contamination half-lives, the cost of running the full ladder for every mechanism. The instrument measures what v0.1 can express. Everything else is labelled pending until a chapter earns it.
So you now have the beginning of one of the book’s two products, and nothing yet worth calling a memory system. That is deliberate. What you have is a way to expose memory failures.
The next chapter builds the strongest simple memory system that could plausibly work, then measures where it breaks. That is the rhythm of the whole book: define, measure, build, fail, diagnose, improve, measure again.
The payoff recurs throughout the rest of the book and reaches its integrated form in the final chapter. Attach a memory system. Define the tasks your application needs memory for. Run them through the instrument. Look at where memory succeeds and where it fails. Change the architecture only when the failure justifies it. Rerun the same instrument. Find out whether the change actually helped.
The book teaches a memory architecture. It also teaches how to measure any memory architecture, long after the last page.