← Memory From First Principles

Can Memory Be Trusted?

Once remembered information can change behaviour, memory becomes an authority channel. This chapter tests what happens when remembered history is stale, untrusted, malicious, private, or deliberately poisoned.

The book’s definition forces this chapter. Memory is when retained past experience changes present behaviour — and Chapters 12 and 13 showed that change running in both directions, including harmful actions under wrong memory and under some fallback conditions. If a line in retained history can change what the agent does today, then admitting that line is no longer harmless. The progression the book has been climbing ends one step further than behaviour:

storage
    ↓
retrieval
    ↓
memory
    ↓
behavioural influence
    ↓
authority

A memory can be highly relevant but untrusted, true but unauthorised for this task, historically accurate but superseded, authoritative but irrelevant, useful but private, well-sourced but maliciously crafted, agent-authored and self-reinforcing, or derived from a source that was later deleted. Retrieval score answers none of these. A high score establishes neither authority nor standing to influence behaviour, and this chapter measures the gap.

A staged gate, not a score

Trust is a staged admission policy with reason codes from the control flow, built on the reconciliation substrate: provenance and revocation registries as visible system state, derivation and refutation as visible lineage, instruction screening as deterministic text rules, corroboration as disjoint source lineages. The policy gives its trust constraints precedence over usefulness — relevant plus malicious is still denied, useful plus cross-scope is still denied, historically true plus superseded is unavailable as current guidance. Maliciousness is never an input flag: the gate must earn poison blocking from content and structure, or fail visibly.

Three simplifications compete with the full gate, pre-registered with a firing match rule: provenance-only deny rules (S1), plus authority and scope (S2), plus corroboration on consequential tasks (S3). A simplification is adopted only if it is breach-free and no worse on task success, harm, attack rate, and retention.

The gate is the chapter’s whole architecture, and its outcomes are categorical rather than graded:

    flowchart TD
    CAND[Candidate memory] --> GATE{Trust and admission gate}
    GATE -->|admit| ADMIT["standing to influence"]
    GATE -->|quarantine| QUAR["contained, still recorded"]
    GATE -->|seek more| SEEK["deferred"]
    GATE -->|reject| REJ["no standing"]
    ADMIT --> BEH[Behaviour]
    SEEK -.->|if support arrives| GATE
    style ADMIT fill:#3978c5,color:#fff
  

Only ADMIT reaches behaviour directly. The other three routes refuse present influence in different ways, and SEEK is the one shown with an explicit path back to the gate if more support arrives.

The four outcomes are not a spectrum. Admission is a categorical decision with three distinct refusal paths, and a memory that is quarantined has not been half-admitted — it has been contained, and recorded as contained.

Fixtures that attack

Eight development and eight held-out evaluation fixtures mix benign units (retention measured) with five adversarial kinds: high relevance plus low authority (losing preferences, ungrounded echoes), high relevance plus poisoned instruction (skip-validation directives, facade deletion), high authority plus stale (superseded decisions and benchmarks), cross-scope but semantically plausible (a second project’s SQLite decision on the same topic), and corroborated but conflicting (refutes-linked pairs with independent support each). Revocation, restriction, unknown source classes and self-reinforcing restatements complete the set. Expected admission is pre-registered per fixture; behaviour is measured, never assumed.

Book result. The gate matches all 8 expected admissions on both splits. Llama primary, memory-matters fixtures:

ConditionTask scoreHarmful tasks
T0 no memory0.3100/7
TU unsafe top-k0.5360/7
TT trust gate0.6790/7
S1 provenance-only0.5360/7
S2 + authority + scope0.5360/7
S3 + corroboration0.5360/7
TO oracle-safe0.6430/7

Frozen runs: trust-dev-v1, trust-eval-v1-llama, trust-eval-v1-muse (19 live calls); grader v2; mapping frozen on dev.

Attack success under the full gate is 0.0 against 0.57–0.86 for the simplifications; benign retention is 1.0 everywhere; no breaches; verdict FULL_EARNS_COMPLEXITY. The development split showed the same shape with teeth: unsafe admission harmed twice (poison-driven backend misconfiguration, facade deletion) where the gate held at zero harm. On two cases the gate also scores above the oracle-safe minimal set (quarantine-plus-context 1.0 against 0.75 on contested evidence). That does not make the gate safer than the oracle; it shows that the minimal safe set is not a behavioural ceiling, because additional SHOULD-grade material can still help the reader.

The strong reader changes the interpretation.

The full gate again matches all 8 expected admissions with zero attack and full retention, and no simplification satisfies the pre-registered replacement rule. But the gate breaches the pre-registered task-success criterion — gated contexts average 0.345 against 0.441 unsafe — so the result is not promoted as a quality win on that reader.

Two priced costs drive the breach.

Quarantining both sides of a contested pair costs decisiveness, 1.0 to 0.0. Containment removes the very content the reader was deciding with.

Enforcing revocation costs utility, 0.833 to 0.333. The revoked benchmark is untrustworthy as a source, and still informative as content.

On the recorded safety metrics, the gate holds on the strong reader — zero harm against one unsafe harm, and zero attack success. The measured cost appears in task utility, where standing to influence and informativeness diverge.

Across the two readers, no simplification satisfies the full replacement criteria. They disagree on the utility cost of the full gate, and those results remain separate rather than being averaged away.

What the numbers mean

Instruction handling does not rely on deterministic screening alone: directives refuted by authoritative current records are denied, corroborated ones admitted, and unverified ones quarantined. Corroboration is structural (disjoint lineages, unrefuted), so self-reinforcing restatements sharing one root never corroborate each other. On these fixtures, quarantine fires on the intended contested cases, and benign retention remains 1.0 on both readers, including grounded echoes and corroborated instructions.

Two findings deserve a plain statement. First, revocation enforcement and quarantine have measurable utility prices on the strong reader, so the trust gate is not a free improvement. Second, the irrelevant-task probe intrudes on both readers with memory admitted (echo succeeds bare, fails framed), preserving Chapter 12’s boundary: on a task that needs no historical evidence, admitting memory can itself change behaviour for the worse.

Verdict. The staged trust gate survives as conditional control: on these fixtures it earns its complexity through attack suppression with benign retention held on both readers, but the Muse task-gate breach prevents any universal quality claim. Quarantine and revocation carry measured utility costs on the stronger reader.

What remains unsolved. Extraction reliability for real histories remains outside this result because the gate consumes fixture-authored lineage; the Chapter 10 problem persists underneath. Timestamp-based staleness independent of validity payloads is also untested. The utility cost has transferred across only two readers, and procedures and causal evaluation beyond fixtures remain on the unearned side of Chapter 16’s boundary. The chapter also leaves open when revoked-but-informative content should be denied standing yet remain visible — quarantine excludes it from influence, and silence has its own failure mode.

The gate now decides what may influence behaviour, priced and bounded to two readers; it says nothing about how this piece fits beside everything else the book earned. That is Chapter 18 — The Remembering System.

Research foundations

The outside literature provides precedents for both persistent memory poisoning and policy-shaped defences.

Srivastava and He (MemoryGraft, 2025, preprint) show poisoned procedure templates persisting through retrieval in agents. That is the direct motivation for the poisoning probe.

The OWASP Agent Memory Guard project (Gudur and Rajkumar, 2026, project-reported) independently converges on staged policy with quarantine — allow, redact, quarantine, block. Its benchmark figures are the project’s own and carry no endorsement here.

Yan and colleagues (CRAG, 2024, preprint) evaluate retrieved documents and trigger corrective actions — the closest prior art to a quality gate.

Moskvoretskii and colleagues (AdaRAGUE, ACL 2025) treat model self-confidence as suspect across dozens of methods. That is the standing reason no trust stage here consults it.

References

  • Saksham Sahai Srivastava and Haoyu He, MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval (2025, preprint).
  • Vaishnavi Gudur and Anshul Rajkumar, OWASP Agent Memory Guard: A Runtime Defense and Open Benchmark for Memory Poisoning in LLM Agents (2026, project-reported; architecture verified against repository).
  • Shi-Qi Yan and colleagues, Corrective Retrieval Augmented Generation (2024, preprint).
  • Viktor Moskvoretskii and colleagues, Adaptive Retrieval without Self-Knowledge? Bringing Uncertainty Back Home (ACL 2025).