Does Better Context Change Behaviour?
The past has been stored, retrieved, selected and assembled. None of it counts as memory until it changes what the system does.
Chapter 1 defined memory behaviourally. Remove the past, rerun the present task, and ask whether behaviour changed — and whether the change was an improvement.
Eleven chapters later, that test has never actually been run.
Chapters 3 to 11 built a pipeline — retrieval, structure, association, routing, lineage, temporal state, open loops, frames, derived consequences. Each stage was measured against ledgers of what the context contains.
Then an assembly experiment under a tight token budget produced an anomaly. At a 768-token budget, the assembled context holds required-evidence recall of only 0.72, yet the reader answers at 0.955 key-claim coverage — better than the full raw context at 0.879.
Three explanations fit. The ledger overstates what the task needs. Or the reader answers without the evidence the book calls required. Or the metric is giving false confidence.
Answer quality alone cannot tell them apart. This chapter builds the instrument that can.
A correct answer is not yet memory
A correct answer is compatible with four states of the world:
- The system used memory well.
- The system already knew the answer.
- The system guessed.
- The system needed only part of what the ledger demands.
The key-claim scorer cannot tell them apart. Neither can citation — a model can emit source identifiers it never reasoned from.
For causal attribution here, the cleanest evidence is interventional. Hold the model, the task, the prompt and the decoding fixed. Change only the memory supplied. Then watch whether the structured behaviour changes, and in which direction.
That gives the chapter its two separated measurements. Behavioural influence asks whether memory changed what the system did. Behavioural utility asks whether the change was better. A system can show influence without improvement, improvement without influence is impossible by definition, and a correct answer under both conditions demonstrates nothing about memory at all.
Tasks that require the past
Nine controlled tasks carry hidden behavioural ledgers, written before any model ran. Each has a frozen grader contract stating what earns full, partial and zero credit, what is forbidden, and what counts as equivalent.
At least half are constructed around project-specific facts not supplied by the present task — the configuration of the baseline store, the reversibility rule that refuses a marginal mechanism, the open loops blocking release, the contracted facade that must survive the corpus_import cleanup.
The tasks also require application, not repetition. Choosing a backend. Holding a chapter. Planning a migration. Each is scored through a tiny deterministic project simulator whose transitions, including which actions count as harmful, are fixed in advance.
The ninth fixture is the negative control: three present-state labels to echo exactly. There, any behaviour change under memory is intrusion, not help.
The ladder
Every task runs a fixed reader at temperature zero, under paired conditions that differ only in supplied memory. The canonical run uses one small model; two transfer readers rerun the same frozen contexts through the same grader, reported separately below and never averaged.
The condition ladder runs:
- No memory — the mandatory floor.
- Full visible history (68 units across three projects) — does undifferentiated past help?
- Project-only history (57 units) — separates foreign-project contamination from unselected volume.
- Strong RAG and frame-conditioned selection — the book’s own pipeline, frozen from Chapter 10.
- Assembled memory at 768 tokens — the practical system, whose construction Chapter 14 reports.
- Auditable oracle — the ledger-defined diagnostic ceiling.
Then three interventions complete it: decisive memory removed, decisive memory restored, and deliberately wrong memory as a positive control for influence. A scrambled control replaces decisive evidence with distractors at matched shape.
Same task, same reader, same prompt. The only free variable is the retained past.
The ladder yields both measurements at once: whether memory changed the action, and whether the change was better.
Book result (frozen run
ch12-20260920T204414Z-behavior, grader v2, prompt v2, 9 tasks, 82 controlled outcomes, 3 repeats each on 3 tasks, plus a 76-outcome second-reader run and 10 time-locked transfer outcomes). Headline means are task-averaged over the seven tasks that run the full B0/B2/B3/B4/BO ladder, repeats averaged within task first — the only comparison in which every condition faces the same tasks:
| Memory condition | What it isolates | Success | Harmful tasks | Mean tokens |
|---|---|---|---|---|
| B0 no memory | empty context | 0.226 | 0/7 | — |
| B1 full history (68 units) | every retained passage | 0.048 | 1/7 | 6242 |
| B1P project-only (57 units) | Memory-project history only | 0.179 | 0/7 | 5495 |
| B2 strong RAG | ranked relevant candidates | 0.393 | 0/7 | 2029 |
| B3 frame-selected | ranked decisive evidence | 0.357 | 0/7 | 1909 |
| B4 assembled (768) | evidence within a budget | 0.488 | 0/7 | 1180 |
| BO oracle-auditable | minimal auditable set | 0.524 | 0/7 | 1329 |
| BA decisive removed* | critical evidence withheld | 0.250 | 1/6 | 1365 |
| BR decisive restored* | critical evidence returned | 0.778 | 0/6 | 1480 |
| BW wrong memory* | misleading past supplied | 0.042 | 2/4 | 120 |
| BS scrambled* | matched-shape distractors | 0.167 | 0/2 | 1015 |
*Starred rows run on intervention subsets (6, 6, 4, 2 tasks), not the matched seven. The six-task remove/restore set is the matched seven minus the echo negative control, which has no decisive memory to remove; corpus-cleanup contributes an additional ablation-only outcome outside the headline means. Harmful-task numerators count every recorded outcome including the ablation-only corpus-cleanup case, so B1 shows 1/7 from that outside-headline case while the matched seven record zero, and likewise for BA/BW. Parse rate is 75 of 82; all 7 parse failures occur outside assembled or oracle contexts — full history (B1), project-only history (B1P), frame-selected without assembly (B3), and decisive-removed (BA) — where decisive evidence is buried in unfiltered volume or withheld.
Frozen contexts are hashed; the runtime never sees a hidden label; determinism is checked by repeats on three headline tasks.
Four things in that table matter more than the headline.
Undifferentiated history hurts — for this reader. Full 68-unit history across three projects scores 0.048 against 0.226 with nothing. Project-only history (57 Memory units, the B1P control) recovers to 0.179 — better than full history, still worse than nothing. Both history conditions sit below no-memory, so the earlier reading that only the full dump fails is false. Removing foreign projects recovers much of the loss, but project-only history still trails no memory, so the residual cannot be blamed on cross-project contamination. For this reader and fixture, even correctly scoped unselected history remains harmful. The strong-reader transfer below changes that interpretation: a substantially stronger reader exploits unfiltered history far better. The Perfect Memory Paradox therefore survives in the narrower form the evidence supports: perfect preservation does not imply useful memory, and the effect of undifferentiated history depends partly on the reader.
Better selection is not yet better behaviour. Strong RAG reaches 0.393, frame selection falls back to 0.357, and assembly recovers to 0.488 at roughly three-fifths the tokens. Frame conditioning, which Chapter 10 showed improves context-selection metrics, does not improve downstream behaviour by itself here for either the primary reader or the second reader (0.524 → 0.452 → 0.643 for the latter). Among the non-oracle matched-ladder conditions, assembly produces the highest success for both of those readers while using substantially less context — roughly 39% fewer tokens than selection on matched tasks. The layers interact non-additively: RAG establishes a strong retrieval baseline, frame-conditioned selection improves evidence properties without improving behaviour by itself, and bounded assembly over the framed evidence recovers a behavioural advantage. The separate restoration condition reaches 0.778 on its six-task intervention subset, but that mean is not directly comparable with the seven-task oracle mean. It is evidence about sensitivity to decisive memories, not evidence that restoration beats the oracle.
Removal and restoration track. Across ablation-applicable tasks, removing decisive memory drops success to 0.250 and restoring it lifts success to 0.778. Three full M→MA→MR traces show the pattern individually: fix-store holds 1.0 across all three repeats, collapses to 0.333 with the postgres record stripped, and returns to 1.0 on restore; cite-rule goes 0.0 → 0.0 → 1.0, the restored decisive unit producing the book’s own rule identifier against two neighbour distractors; corpus-cleanup goes 0.833 → 0.0 → 0.833, the ablated run updating docs immediately and deleting the contracted facade. The reader’s behaviour depends on the specific remembered item, in both directions.
Wrong memory is influential and destructive. BW averages 0.042 with two harmful tasks out of four: the bare facade-deletion listing deletes the contracted facade (B6), and the stale SQLite record drives a SQLite configuration (B6 under grader v2, which records acting on superseded operational state as a harmful event). Full history also deletes the facade once. Most strikingly, removing decisive caller evidence (BA) lets the trap win too: the ablated cleanup run deletes the facade with B6 recorded. On fix-store the ladder is progressive rather than divergent: no-memory emits an unusable unknown backend (0.333), strong RAG and frame selection configure PostgreSQL partially (0.833 each), assembled and restored contexts configure it fully (1.0) — while deliberately wrong memory configures SQLite with B6 recorded. The failure-avoidance mechanism Chapter 8 built now has behavioural consequences, but in this run they appear at the wrong-memory extreme rather than between RAG and framing. The two directions together are the chapter’s result in miniature: correct memory can improve behaviour, and wrong memory can also control behaviour. A memory layer that decides what the model may see can cause the harmful action, not merely fail to prevent it. That is the positive control doing its job — the premise Chapter 13 has to answer.
Dimensions, not one number
Task success averages hide where memory helps. Reported separately, and task-averaged over all tasks:
- Constraint adherence moves 0.083 with no memory, to 0.500 under RAG and selection, to 0.583 assembled, to 0.600 at the oracle. Removal reverses it to 0.250; restoration completes it at 0.750.
- Open-work continuation moves 0.000 to 0.625 assembled, 0.667 at the oracle — with restoration reaching 1.000 from 0.000 removed.
- Failure avoidance is 1.000 everywhere, except under project-only history where a parse failure zeroes the single task, and under wrong memory where the stale record wins.
- Goal adherence moves 0.333 to 0.500 to 0.429, with restoration lifting it to 0.714.
Harmful actions are zero under no-memory, selection, assembly and oracle. They are nonzero under full history, decisive-removed cleanup, and wrong memory.
So the pipeline does not raise every dimension together. Constraint adherence and open-work continuation improve under some memory conditions, failure avoidance is mostly saturated except in the specified failure cases, and goal adherence remains non-monotonic.
The negative control marks the boundary exactly. On the echo task, no-memory scores 1.0, while RAG and frame-selected context both go silent — abstaining on the grounds that the labels are not found in the evidence.
The negative control shows that memory can hurt a task that requires no historical evidence. With evidence present, the reader treats the memory store as the source of truth and withholds the trivially correct echo.
The strict influence table over 73 paired comparisons reads: 14 full-success contributions, 2 harmful, 0 memory-not-necessary, 57 failed-to-repair.
That last number needs the directional table beside it, because the strict table credits only perfect treated scores.
Across the same 73 pairs: structured actions change on all 73, scores improve on 32, degrade on 8, and stay equal on 33.
A 0.00 → 0.75 move counts as failed-to-repair in the strict table, and as improvement in the directional one. Both are published, each labelled for what it is.
The honest majority is partial movement, not full success. Across these paired comparisons, changing the supplied memory changes the structured action every time; scores improve on 32 pairs and degrade on 8.
The 0.72 / 0.955 paradox, resolved
The anomaly that motivated this chapter dissolves into two findings, one about the ledger and one about the reader.
First, the ledger overstates unit-level necessity on some tasks.
On the architecture task, four of six MUST units are absent, yet coverage is 1.0. The reason is a substitution. An open-loop record, unlabelled in that task’s ledger and therefore a distractor by default, carries the same actionable content as the missing MUST unit — plus SHOULD-grade support for the remaining claims.
The same substitution pattern covers the other high-coverage, low-recall tasks.
This chapter confirms it behaviourally. Under the identical 0.333-recall context, the reader holds the chapter for exactly that reason, scoring 1.0.
So the evidence was sufficient for the behaviour. The ledger was conservative about which units count.
That is recorded as a ledger finding for the next benchmark version — alternative support between the loop record and the experiment-pending unit. Not as a silent label edit, which the version rules forbid.
Second, the reader sometimes answers well from evidence that the ledger scores as incomplete. At 768 tokens the composed policy protects disagreement and licences at the cost of required recall (0.72), yet answers better than the full context. That coexistence does not identify a single cause: substitutions, partial evidence, and reader capability can all matter. Ledger-defined evidence sufficiency and what a particular reader can do with the supplied context are different quantities. The 0.72 recall figure therefore cannot be read as proof that the missing evidence was unnecessary in general.
Two readers, one direction
A second reader shows the same direction with different magnitudes on matched tasks: 0.000 with no memory, 0.524 under RAG, 0.452 selected, 0.643 assembled, against an oracle of 0.536.
Three reader-dependence findings stand out.
It starts from zero. The smaller reader abstains without memory almost everywhere, so memory’s measured delta is larger even where its ceiling matches.
Its failures are more sensitive to supplied context. Under wrong memory it follows the stale record, ships the unready chapter, and deletes the contracted facade twice over. It also deletes the facade with no memory at all, where the primary reader holds. The comparison therefore shows influence in both directions: supplied memory can trigger harmful actions, but it can also supply a constraint the reader does not otherwise follow.
The ledger oracle is not a behavioural oracle for it. Assembled context (0.643) beats the oracle (0.536). That result shows that the ledger-minimum set is not the context that maximises this reader’s behavioural score; it does not by itself identify which omitted material explains the gap.
So the instrument measures the memory–reader interaction, not memory alone. The architecture supplies better evidence, but readers differ in floors, ceilings and suggestibility. The readers are reported separately and never averaged.
Does a stronger model need memory?
Perhaps the memory system only helps because the development reader is small. The transfer wave exists to test that objection rather than argue against it: the same frozen contexts, grader, prompt and fixtures through muse-spark-1.3-contributor (reasoning low, temperature zero), with the reader as the only variable. An early single-sample transfer was then confirmed with three repeats per task per condition, a session-isolation audit, and a token-matched remove/restore control (run identifiers captioned below). The confirmatory matched means:
| Condition | Primary reader | Ministral | Muse Spark 1.3 |
|---|---|---|---|
No memory (B0) | 0.226 | 0.000 | 0.119 |
Strong RAG (B2) | 0.393 | 0.524 | 0.413 |
Frame-selected (B3) | 0.357 | 0.452 | 0.786 |
Assembled (B4) | 0.488 | 0.643 | 0.667 |
Oracle (BO) | 0.524 | 0.536 | 0.726 |
Decisive memory removed (BA) | 0.250 | 0.417 | 0.472 |
Decisive memory restored (BR) | 0.778 | 0.528 | 0.736 |
Wrong memory (BW) | 0.042 | 0.000 | 0.167 |
Muse without memory (0.119) does not exceed the small reader without memory (0.226): the stronger model abstains honestly where the small one guesses, rather than solving memory-dependent tasks from pretraining. Structured memory still improves it (B4 0.667, residual benefit +0.55), decisive removal still degrades it and restoration still recovers it (BR − BA +0.26), and deliberately wrong memory still causes a harmful action — the stale SQLite record drives a SQLite configuration with B6 recorded, while other wrong-memory tasks abstain safely. Capability and suggestibility coexist.
On tasks whose decisive facts are arbitrary project history — which backend was chosen, which loops block release, which rule governs restatement — Muse without memory scores 0.056 and with assembled memory 0.611. On these fixtures, the stronger reader does not recover those arbitrary historical facts from the present task alone; assembled project memory supplies information its general capability does not.
The restore effect is not explained by token volume alone. A token-matched control (decisive removed, nondecisive context restored at matched token cost) does not recover the restoration effect on either reader (see caption): adding nondecisive tokens back leaves the mean near removal, while returning the decisive item raises it substantially. One exception is reported, not averaged away: on corpus-cleanup, generic support partially recovers behaviour, so nondecisive context contributes there while decisive content still adds the remainder.
Two reader-dependence findings qualify the chapter.
Undifferentiated history is reader-dependent. The strong reader scores 0.571 on full and project-only history, against the small reader’s 0.048 and 0.179. Unlike the primary reader, it scores well above its own no-memory floor of 0.119 with unfiltered history. Structured conditions still score higher — 0.786 frame-selected, 0.667 assembled, 0.726 at the oracle — but unselected history is not behaviourally harmful relative to no memory for this reader. What transfers is the narrower conclusion: preservation alone does not guarantee the best behaviour.
Selection beats assembly for the strong reader. Frame-selected context (0.786) beats assembled context (0.667), reversing both local readers.
That reversal is task-specific. One task cites correctly under selection in all repeats, but abstains under assembly in two of three. A structural diff shows the decisive rule text present identically in both, so loss of that decisive text does not explain the reversal. The remaining difference belongs to the interaction between the reader and the rest of the assembled context.
Assembly primarily buys efficiency for the stronger reader, with that exception recorded.
The distinction the wave earns is threefold: model capability (what the reader can reason about), project memory (what historically happened in this particular project), and memory policy (which historical state may influence current action). Across readers with very different no-memory behaviour, structured project memory continued to alter downstream actions: the stronger reader did not eliminate the effect.
Book result (transfer runs captioned below; contexts hash-identical to canonical throughout). The measured memory effect survived a substantially stronger reader: Muse used structured project memory far more effectively than no memory on the controlled tasks, including tasks whose decisive information was arbitrary project history. Removing decisive memories still degraded behaviour, restoring them recovered behaviour, and deliberately wrong memory remained capable of causing harmful action. This establishes reader transfer across the tested models and fixtures, not a universal law over future models.
Transfer runs: sr1-ch12-20260920T230129Z-muse (exploratory), sr1c-ch12-confirm-20260920T234546Z-muse (confirmatory, 3 repeats), token-control runs ch12-token-control-*, session audit sr1b-session-audit-*; token-matched control means across removal / token-matched / restored (BA/BT/BR): llama 0.214/0.167/0.643, Muse 0.548/0.552/0.853.
Real-project transfer
Five time-locked repository questions run at the frozen commit, manually adjudicated, and scored separately from the fixtures. Memory-supplied answers average 0.95, against 0.35 without.
Per task: run-report 0.75/0.25, fair-gap 1.0/0.0, promote-decision 1.0/1.0, frame-remedy 1.0/0.5, null-rule 1.0/0.0.
Two of those need annotation.
The promote-decision task is non-diagnostic. Its question states the breach, so rejection needs no memory, and both conditions score 1.0.
The null-rule task works as designed. Copying the demonstration scores 0 without memory; stating explicit nulls scores 1.0 with it.
At small n, the real-project transfer points in the same direction as the controlled comparison, with every disagreement preserved in the frozen file.
What this chapter earns, and what it does not
Against the outcomes declared before the run, this is a Type A result on fixture evidence. On the matched seven tasks, assembled memory improves success over no memory (0.226 → 0.488) and strong RAG (0.393 → 0.488). On the separate six-task intervention subset, removal drops success to 0.250 and restoration lifts it to 0.778, with attribution traces naming the decisive items. For the primary reader, full history also scores below no memory (0.048 against 0.226), showing that undifferentiated history can be harmful; the stronger-reader transfer above shows that this last effect is not universal across readers.
Five demotion clauses apply, and each materially limits the claim.
The fixtures are synthetic and few — 9 controlled, 5 transfer.
The primary reader is one small model. Two transfer readers check reader-dependence; they do not replace it.
One fixture is non-diagnostic. The prose-review task scores 0.0 under every condition, including the oracle. It does not separate conditions, and is kept as a published instrument failure rather than dropped.
One task is behaviourally flat. The credit-rule task also scores 0.0 everywhere including the oracle — without action demonstrations the reader abstains or omits the mechanism name. It measures format compliance rather than memory use. Kept, not repaired after the fact.
Repeats cover only three tasks. Determinism is confirmed there; the remaining outcomes are single-shot at temperature zero.
Nothing here claims that memory improves behaviour in general. The claim is narrower: on memory-dependent controlled tasks, this reader acts better with structured project memory — and targeted interventions attribute part of that improvement to specific remembered evidence.
What remains unsolved. The instrument now exists, and the first thing it reveals is how much behaviour it cannot yet explain: 57 of 73 pairs never reach full success under either condition, and wrong memory averages 0.042 while producing harmful actions in two cases. On these fixtures, Chapter 12 establishes that supplied memory can improve behaviour, that removing specific memories can reverse gains, and that stale or inappropriate memory can also drive harmful action. Once memory has behavioural power, the next question is no longer merely which memories are relevant. It is when the system should trust its own framing strongly enough to let those memories influence behaviour. That is Chapter 13 — When the Frame Is Wrong: whether memory can be made safer without being made weaker — the abstention path Chapter 10 specified but never built — and it now has a way to be measured.
Research foundations
The outside literature converges on the instrument rather than the mechanism.
Mem2ActBench (Shen and colleagues, ACL 2026) draws the operative cut between remembering information and using memory to act, with tool tasks reverse-generated so the past is necessary by construction. This chapter’s fixtures follow the same rule in project form.
MemoryArena (He and colleagues, 2026, preprint) couples acquisition with later use across interdependent sessions, and reports that saturated recall collapses in agentic settings — the same direction as the divergence above, approached from the opposite side.
CUB (Hagström and colleagues, ACL 2026) supplies the warning this chapter heeds throughout. Context present does not imply context used, and simple synthetic wins inflate.
DRUID (Hagström and colleagues, ACL 2025) motivates the real-project transfer: synthetic utilisation results overstate, so a small adjudicated set anchors the fixtures.
LOCOMO-CONV (Chang and Chen, 2026, preprint) names the discrepancy class this chapter classifies as substitution — strong retrieval not translating into response quality, with silent grounding where memory helps without surfacing the gold fact.
Agent Workflow Memory (Wang and colleagues, ICML 2025) is precedent: remembered structure can move task success. MemBench (Tan and colleagues, ACL Findings 2025) is background: effectiveness without a decision metric is not behaviour.
References
- Meta AI Research, Introducing Muse Spark 1.3 (2026). Primary model announcement for the strong-reader transfer reader; vendor-authored.
- Yiting Shen and colleagues, Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents (ACL 2026).
- Zexue He and colleagues, MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks (2026, preprint).
- Lovisa Hagström and colleagues, CUB: Benchmarking Context Utilisation Techniques for Language Models (ACL 2026).
- Lovisa Hagström and colleagues, A Reality Check on Context Utilisation for Retrieval-Augmented Generation (ACL 2025).
- Wen-Yu Chang and Yun-Nung Chen, When Users Don’t Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents (2026, preprint).
- Zora Zhiruo Wang and colleagues, Agent Workflow Memory (ICML 2025).
- Haoran Tan and colleagues, MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents (ACL Findings 2025).