← Memory From First Principles

When the Frame Is Wrong

Frame establishment now rests on explicit evidence relations, not classifier confidence — and survives a simplification test across two readers.

Chapter 12 ended with leverage. In the wrong-memory condition, success averages 0.042 across four tasks, with harmful actions on two.

In that controlled intervention, a layer deciding what the model may see can cause a harmful action — not merely fail to prevent one.

So the next question is when the system should trust its own framing that strongly.

Frame establishment here does not rest on classifier confidence. It is derived from explicit, temporal, provenance-preserving evidence relations. A reconciliation records that two representations relate, without ever asserting what is true.

Reconciliation first, classification never

A reconciliation records that an intent and a commit describe one event, that a commit closes an obligation, or that two sessions merely correspond. It also records when the relating was asserted, who asserted it, and whether it claims to settle anything or only notes a correspondence.

Resolving one classifies the relationship at a date as corroborated, stale, conflicting, invalidated, cross-scope, or correspondence-only.

It never answers whether the records’ content is true. The implementation enforces that structurally:

  • Correspondence cannot resolve.
  • Contradictions refuse rather than merge.
  • Cross-project relating quarantines without explicit licence.
  • Later revocation invalidates.
  • Ended validity goes stale.

Verification and validity paths are untouched. A dedicated regression test pins the claim that outcomes are identical with and without reconciliations.

Frame establishment is then derived deterministically from those resolutions — corroborated, weak, stale, conflicting, or unknown hints per frame subject, each naming its basis.

The policy maps them:

  • Corroborated → hard framing.
  • Weak → soft framing, without exclusion.
  • Stale or conflicting → query-only fallback.
  • Unknown → fallback, except where abstention is acceptable (decline) or consequential work needs memory (request more evidence).

Correspondence-only evidence routes to fallback action explicitly. Topical but unestablished, it licenses acting without the frame — never refusal, and never hard framing.

No scalar confidence appears anywhere. Reason codes arise from the control flow actually taken.

The fixtures and the simpler alternatives

Twelve development and nine held-out fixtures share small ledgers and frozen behavioural rows across six reconciliation cases — same-event current, correspondence-only, stale, conflicting, invalidated, cross-scope — so the unit under test is the decision procedure alone. The mapping stood as authored on development and was frozen before evaluation opened.

Two simplification competitors were pre-registered before evaluation. S1 frames hard wherever evidence is corroborated and falls back otherwise; S2 always soft-frames. The match rule compares harm, oracle-match (agreement with the ledger-oracle framing on each scenario), and benefit gap. If either matches the full gate, the chapter adopts it. That rule determines how the result is interpreted below.

Book result. The gate matches all 9 held-out evaluation classes on both readers with no breaches. Llama primary:

ScenarioGate actionF6F1FOS1S2
arch-hardHARD1.01.01.01.00.75
fix-softSOFT1.00.8331.00.8331.0
arch-staleQUERY_ONLY1.01.01.01.01.0
cite-confQUERY_ONLY0.250.250.250.250.25
ship-reqREQUEST0.00.1670.1670.1670.167
proseQUERY_ONLY0.00.00.00.00.0
fix-corrQUERY_ONLY0.8330.8330.8330.8331.0
fix-invalQUERY_ONLY0.8330.8330.8330.8331.0
irrelQUERY_ONLY0.00.0—0.00.0

Frozen runs: ch13s-dev-v1, ch13s-eval-v1-llama, ch13s-eval-v1-muse; grader v2; mapping frozen on dev.

Oracle-match 8/9 (the miss is the priced ship-request abstention), no harm, benefit gap 0.0. S1 matches the gate on every pre-registered term — same oracle count with a different single miss, no harm, gap 0.0 — while always-soft fails the benefit gap (0.25, losing corroborated-hard benefit). On the primary reader the verdict fires as pre-registered: adopt S1.

On the strong reader, that simplification fails.

The full gate again matches all 9/9 gate classes with no breaches. But fallback is harmful where it was benign for the primary reader — three fallback rows score 0.0 with harm recorded.

The fallback-only simplification, which falls back everywhere uncorroborated, carries all three of those harms and fails the match.

Meanwhile always-soft, which failed on the primary reader, matches the gate here instead: no harm, oracle 7/9, gap 0.0. Soft framing rescues exactly the rows where this reader’s fallback harms.

Taken separately, each reader makes a different simplification satisfy the pre-registered match rule.

Both per-run verdicts are reported as computed, and neither is adopted.

Because one frozen policy must serve both readers without per-reader tuning, neither simplification transfers across them. Fallback-only harms on the strong reader. Always-soft loses benefit on the primary reader. The full gate is the only tested policy with no breaches on either.

The simplification test therefore matters in both directions — it rejects always-soft on one reader and fallback-only on the other.

Across these two readers, the extra policy branches matter in the reader-specific failure cases rather than as a broad mean-score gain.

The resulting rule is simple to state. Corroborated evidence goes hard and retains benefit on both readers. Weak evidence goes soft without exclusion. Stale or conflicting evidence falls back to query-only, as does correspondence-only evidence. Reconciliation never promotes mere correspondence into truth.

Costs, carried over and repriced

Requesting more evidence costs the ship task on both readers — 0.0 against an answering oracle at 0.167. That is the priced abstention already visible in the 8/9 oracle-match result, now reproduced across both readers.

Correspondence and invalidation fall back and act well on the primary reader (0.833, oracle-matched). The same fallback harms on the strong reader for both of those rows, oracle included on one.

The result is narrower. On these fixtures, fallback is not inherently safe: the strong reader records harm on three fallback rows where the primary reader does not.

That reframes the earlier over-conditioning finding rather than repeating it: the danger moved from refusal to action.

Broaden-retrieval stays unmapped — the earlier rejection stands, and no fixture forces it. The prose floor, 0.0 everywhere including the oracle, is kept as a published instrument failure.

What remains unsolved. The cross-reader result supports the full gate only within two readers: each simplification fails on a different one, so broader transfer remains unestablished. Broaden-retrieval remains unmapped, and poison admission remains a separate trust problem. The gate now decides when established framing evidence may influence behaviour; it does not decide how to fit the resulting evidence into a bounded context. That is Chapter 14 — Context Is a Bottleneck.

Research foundations

The outside literature provides precedents for several parts of this control surface: model self-confidence can be unreliable (Moskvoretskii and colleagues; Min and colleagues), abstention can be treated as a control action under cost asymmetry (Huang and colleagues), corrective gates and complexity-based routing precede the frame gate (Yan and colleagues; Jeong and colleagues; Jiang and colleagues), and poisoned experience can persist through retrieval (Srivastava and He).

References

  • Viktor Moskvoretskii and colleagues, Adaptive Retrieval without Self-Knowledge? Bringing Uncertainty Back Home (ACL 2025).
  • Dehai Min and colleagues, QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation (ACL Findings 2026).
  • Zhiqi Huang and colleagues, Confidence-Based Response Abstinence: Improving LLM Trustworthiness via Activation-Based Uncertainty Estimation (UncertaiNLP 2025).
  • Shi-Qi Yan and colleagues, Corrective Retrieval Augmented Generation (2024, preprint).
  • Soyeong Jeong and colleagues, Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity (NAACL 2024).
  • Zhengbao Jiang and colleagues, Active Retrieval Augmented Generation (EMNLP 2023).
  • Saksham Sahai Srivastava and Haoyu He, MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval (2025, preprint).