← Memory From First Principles

The Memory Nexus

Once a system has several ways to remember, something has to choose between them. This chapter builds that control layer, measures it against simpler alternatives, and lets the evidence decide whether it earns its place.

Chapters 3 through 5 leave the system with an embarrassment of options. Chapter 3 built hybrid retrieval over raw history with reranking. Chapter 4 added a persistent derived graph with several query modes. Chapter 5 added cue-conditioned associative propagation over that graph. Each chapter earned its mechanism conditionally, and each left the cheaper layers available underneath. The question none of them answers is the one a deployed system meets first:

Once an AI has several legitimate ways to access memory, what decides which memory process should run for the situation it faces now?

This chapter builds a candidate answer, measures it against simpler alternatives, and lets the evidence determine whether a separate memory-control layer is justified.

What Chapter 5 actually showed

Chapter 5’s frozen runs are retrieval-level on a synthetic 73-node graph: cue-conditioned propagation modestly improved recall at a matched budget, constraints mattered more than the propagation rule, and branch separation emerged at seeding rather than during the walk. That narrow result motivates the control layer without proving it works:

Book hypothesis. The next architectural problem is choosing the right entry point and memory strategy, not propagating more aggressively.

Every routing claim below states the representation and mechanism scope it was tested under.

The Memory Nexus

The working definition is deliberately narrow:

The Memory Nexus is a policy over memory resources.

It is not the memory store, the knowledge graph, the associative graph, the final reasoning model, or the user’s agent. Its job is to decide, given the current situation and the available memory mechanisms, what memory operation to perform next:

current task/query
      ↓
observable state
      ↓
Memory Nexus
      ↓
memory action
      ↓
observation/result
      ↓
stop or choose another memory action

A memory action names a capability plus the budgets it runs under, because using the graph with a three-hop limit and provenance required is a different decision from using the graph unbounded. The registry contains the mechanisms built in Chapters 3–5, plus control and fallback conditions: NONE, RAG, GRAPH_BASIC, GRAPH_LOCAL, GRAPH_GLOBAL, GRAPH_DRIFT, ASSOCIATIVE, and RAW_EVIDENCE. Later chapters may register temporal resolution, belief resolution, or consolidated memory without rewriting the router. Those mechanisms are not implemented here.

Staging matters. This chapter routes memory mechanisms with the reader held fixed. Budgets remain part of each memory action, but reader and model selection do not vary with the route, so a stronger reader cannot masquerade as a better memory policy.

The homunculus constraint

A controller that already knows the answer has not routed; it has answered. The Nexus therefore sees only a fixed observable state assembled from the query, capability state, budgets, and prior-action diagnostics, never from the evaluator’s ledger:

query text and embedding-length features
interrogative form and lexical signals
linked entities from the graph snapshot
available and degraded capabilities
context and latency budgets
prior actions and their diagnostics
confidence, contradiction, and disagreement signals

Ledger fields such as task family, expected sources, and expected state are refused by construction. A leakage audit runs in the test suite and in the experiment driver: any state object carrying an evaluator-only field is rejected. The one exception is a labelled control policy that routes on the ledger’s task family explicitly, so the chapter can measure exactly what a leaked label is worth. Its score is an upper bound on query-classification routing, never a result the Nexus may claim.

How routing appears elsewhere

The literature review behind this chapter is kept in the project’s research notes rather than reproduced here. Five lines of prior work shaped the build, each verified against its landing page before entering the bibliography.

Model routing establishes that routing can be worth building. Three systems are worth naming:

  • RouteLLM learns quality/cost routing between strong and weak models from preference data, with thresholds tuned to explicit quality targets.
  • Router-R1 treats multi-model routing as a sequential decision process trained by reinforcement, with think and route actions and an explicit cost reward.
  • RouteLMT frames routing as budget allocation, and identifies the operative signal as expected marginal gain — how much the large model improves on the small one. Difficulty and absolute-quality proxies both misallocate budget.

The transferable claim is modest. Routing between capabilities under cost pressure is an established engineering problem with known solution shapes. This book tests whether those shapes fit memory.

Mixture-of-Experts contributes mechanism warnings. Sparsely-gated layers demonstrate that learned gating works and that it collapses without explicit anti-collapse machinery. Switch Transformers demonstrate that the simplest viable router is often sufficient. The Nexus adopts the simplicity bias — fixed and rule policies are the bar every learned router must clear — and reports route counts so that collapse toward one capability would be visible rather than hidden.

Cognitive gating is treated as hypothesis generator, not precedent. PBWM-style models demonstrate that input and output gating policies for working memory can be learned from reinforcement signals, and supply the vocabulary of selective updating that this chapter borrows. No software module here corresponds to any brain structure, and the chapter states that plainly: the useful observation is computational, namely that selective access under learned control is a coherent design, not that the architecture recreates anything neural.

Bounded decision models supply the chapter’s sharpest architectural question: whether the Nexus needs an autoregressive language model at all. Vendor offerings in this space make strong claims about calibrated structured choice from non-autoregressive models, but those claims arrive as product announcements with undisclosed internals, and the book treats them as vendor-reported rather than established. The question is tested with the book’s own small classifier instead of a vendor API, and no book component borrows vendor terminology.

Resource-rational control supplies the theoretical frame. Metareasoning treats computations as actions with expected value derived from their effect on the next physical decision, which grounds both the marginal-gain escalation rule and the stopping rule: continue only while another memory operation is worth its cost. The same literature is honest that exact metareasoning is infeasible, so the chapter tests greedy and hand-authored approximations instead of claiming optimality.

Implementation

The Nexus coordinates the capabilities from Chapters 3–5 without replacing them:

query
 ↓
Nexus (capability registry + policy + controller)
 ├─ Chapter 3 RAG and raw-evidence fallback
 ├─ Chapter 4 graph modes (basic, local, global, drift)
 └─ Chapter 5 associative retrieval over the real graph snapshot
 ↓
shared fixed reader
 ↓
Chapter 2 Measurement Instrument (unmodified scorers)

Each capability contributes only its own evidence under the chosen budget, so the alternatives stay separable. A capability that always contained RAG could never lose to RAG, and the comparison would be vacuous.

Availability is checked, never assumed. A capability whose backing system is missing is never registered. One whose health check fails is registered and marked degraded, so routing can see it and decline it.

One more separation matters. Control state — policies, versions, thresholds — is kept apart from memory state — graph contents, sources, claims. Retraining a router must never look like the project changing its mind.

Eight policies are implemented. FixedPolicy and RandomPolicy are the floors. RulePolicy matches on query features. LLMPolicy makes a JSON-bounded structured choice from a frozen model and prompt. ClassifierPolicy is a multinomial logistic regression over dense state features, trained from scratch with no new dependencies. SequentialPolicy goes cheap first and escalates on diagnostics. OraclePolicy and CheapestAdequateOracle are the ceilings.

The sequential controller enforces three independent loop guards — maximum action count, cost ceiling, and no repeated capabilities — each tested as the binding constraint.

Every decision emits a trace recording the observable state, candidates, scores, confidence, and policy version.

One boundary holds here as everywhere else in the book. A trace explains why a mechanism was chosen. It never explains why the resulting claim is true. Factual provenance still bottoms out in source evidence.

Utility before routing

A router judged on answer quality alone would invoke every mechanism on every query and be useless; judged on cost alone it would retrieve nothing. The chapter reports raw dimensions first — task quality, evidence quality, harm, latency, tokens, model calls, graph expansions — and applies a scalar only where an ordering is required (oracle selection, regret). The weights are a stated choice recorded in every run manifest, not a discovered truth:

$$ u \;=\; \text{quality} \;+\; 0.5\,\text{evidence} \;-\; 0.01\,\text{cost} \;-\; 0.5\,\text{harm} $$

with an adequacy threshold of $0.75$ (nexus-utility-v0.1).

Quality is the mean of the instrument’s own task metrics for the cell. Harm is the unsupported-source rate. No scorer was added to the instrument for this chapter; routing-specific quantities (choice, regret, under/over-routing, action counts) are computed from the matrix, not from new judges.

Establish specialisation first

Book result. The task-by-capability matrix is frozen as run ch6-20260919-nexus: 8 of the 20 routing tasks measured, 46 cells across 7 capabilities, one shared reader, matched context budgets. DRIFT is unmeasured and 12 tasks are pending live runs; nothing below is stated beyond those 8 tasks.

Mean fixed-policy scores over each capability’s measured cells (coverage is uneven at 6–7 tasks per capability):

capabilityqualityevidenceharmcostlatencyutility
RAW_EVIDENCE1.0000.5660.0002.509.3s1.258
ASSOCIATIVE0.9290.5740.0001.2014.4s1.203
RAG0.9290.5000.1431.0016.4s1.097
GRAPH_LOCAL0.9170.3470.4253.0064.7s0.848
GRAPH_BASIC0.9290.2980.7861.5055.5s0.670
GRAPH_GLOBAL0.7780.3470.76712.0069.4s0.448
NONE0.1430.1430.0000.003.5s0.214

Within the measured cells, no single capability wins every task. Unique wins spread across five capabilities: associative retrieval on four tasks, graph-local, no-memory, RAG, and raw evidence on one each. Every measured task separates at least two capabilities on the reported utility or component dimensions, although only one separates them on task quality, as discussed below. The falsification condition — one mechanism winning everything — did not occur on these 8 tasks.

Two findings inside that table deserve emphasis, because both cut against the obvious story.

The graph modes are the harm centre on this measured slice. They carry substantial unsupported-source harm (0.43–0.79), while raw evidence, associative retrieval, and no-memory carry none. These graph query modes are not merely expensive here; they also introduce measured harm. That is exactly why the raw-evidence fallback exists.

Only one task separates on quality. On the provenance question, raw evidence reaches 1.0 while every other measured capability scores 0.5. One discriminating task in eight is a thin basis for architecture, and the chapter does not pretend otherwise.

The oracle Nexus

Book result. Perfect per-task routing over the 8 measured tasks reaches quality 1.000 at mean cost 1.41 and utility 1.325. Against the best fixed policy, quality headroom is 0.000 and utility headroom is 0.067 on this 8-task fixture slice (12 tasks and DRIFT unmeasured; see below).

On this measured slice, the oracle’s headroom is entirely cost-shaped: the cheapest adequate route per task holds quality at 1.000 while cutting mean cost to 1.09, routing five tasks to RAG and one each to associative retrieval, no-memory, and raw evidence. The experiment shows no measured quality gain from routing here; it shows a cost gain from declining expensive machinery where cheap machinery suffices. That maps to pre-registered outcome Type C: the Nexus as efficiency layer, not capability unlock.

Rules before learning

Book result. The hand-authored rule policy scores mean quality 0.929 with quality regret 0.071 across 7 scored tasks: six routes to RAG, one to no-memory, one under-route, zero over-routes.

The single under-route is that discriminating provenance task. The rules choose RAG (quality 0.5); the matrix says raw evidence (quality 1.0).

The query features available to this rule policy do not identify that wide unranked evidence is the better route on this provenance task. That miss is informative: it motivates testing post-retrieval diagnostics as routing signals rather than assuming query classification alone can predict fallback value.

The leaked-label control sharpens the same point. Routing on the evaluator’s own task family scores no better than the rules. Knowing what kind of question this is, in the evaluator’s private vocabulary, buys nothing on these tasks — which should temper any enthusiasm for query-classification routing.

The generative LLM router is implemented, frozen, and schema-guarded, but unmeasured: no scored cells exist for it in the frozen suite. It is reported as prototype, not as evidence.

A faster decision engine

The bounded logistic classifier, trained offline on 7 held-in matrix samples with whole-family holdout, recovers oracle quality (1.000) at mean cost 1.21, with zero under-routes and one over-route.

The result is encouraging but fragile. Seven training samples cannot support a generalisation claim.

The narrower conclusion is that, on this fixture, the routing labels can be expressed by a small deterministic classifier. Nothing in the measured gap establishes that generative interpretation of the question is required.

Observe, then choose again

The sequential controller is built, unit-tested, and demonstrable. Live sequential episodes against the real systems are not yet in the frozen suite, so the chapter claims no sequential result.

The motivation for sequential control is specific: the one-shot rule miss occurs where the available query features fail to identify the raw-evidence fallback. A route-observe-route loop could use diagnostics from the first memory action before deciding whether to stop or escalate. That remains a design hypothesis until live sequential episodes are measured.

The stop-policy analysis is implemented against the matrix and runs as soon as sequential episodes exist. Stopping too early is under-routing. Stopping too late is over-routing. The suite measures both rather than asserting a threshold.

Under-routing, over-routing, and regret

Routing regret — oracle utility minus chosen-route utility, split into quality regret and cost regret — is the chapter’s primary control metric. It separates two failures that accuracy alone conflates: choosing too little machinery (the provenance task routed to RAG) and invoking expensive machinery unnecessarily (graph-global synthesis where RAG already answers). Pareto reporting accompanies every comparison: a router is useful when it holds near-best quality while avoiding needless expensive operations; neither cheap-at-collapsed-quality nor accurate-at-everything-always counts as victory.

Marginal gain, adopted from the routing literature, is the escalation hypothesis: ask not whether the query is hard but what improvement is expected from invoking the next capability over the current one. A linguistically complex question can be easy for RAG; a simple one can need the graph. The chapter implements the framing in its utility and its cheapest-adequate oracle. Gain prediction itself — in the RouteLMT sense of probing internal representations — is reserved work: the Nexus stays at the capability level and assumes no access to reader internals.

When the Nexus itself fails

Router failures name control faults. They are kept separate from Chapter 2’s memory-and-reasoning failure classes, so that bad routing with a lucky answer stays distinguishable from good routing with a bad reader.

The taxonomy — wrong capability, under-routing, over-routing, bad seed, premature stop, failed or excessive escalation, loops, cost overrun, confidence misread, feature leakage — is implemented in the controller and exercised by tests.

Two of those classes already have measured instances: under-routing on the provenance task, and cost-bearing over-routing by the graph modes.

One problem is flagged without being solved. A router that mostly chooses RAG collects evidence mostly about RAG, and naive retraining on that log would entrench the majority route. Counterfactual and off-policy evaluation are reserved for later.

Did a control layer earn its place?

Against the pre-registered outcome types, the 8-task evidence maps to Type C with the Type B shadow close behind — Type B being the outcome where fixed policies already suffice and no control layer is needed. Within the incomplete measured matrix, no fixed mechanism wins every task, so a routing problem exists in the weak sense; but the oracle’s headroom is cost-only, the rule policy captures most of the available utility at zero learning cost, and the classifier result is too small to trust. The evidence therefore supports the Nexus only as a conditional efficiency and safety layer: declining expensive derived memory where cheap memory suffices and retaining fallback where derived memory harms.

Then the materiality test applies. A utility gain of 0.067 over the best fixed policy, on 8 tasks, with 12 tasks and DRIFT unmeasured, does not justify control infrastructure in production.

The remaining question is whether the pattern survives broader measurement. The current matrix shows distributed wins, harm concentrated in graph modes, and only one task that discriminates on quality — enough to keep the control hypothesis open, not enough to settle it.

So the architecture stays conditional, and the two ways it could resolve are already named. If the full task matrix shows one fixed mechanism dominating, the Nexus shrinks to a cost guard, or disappears. If sequential control captures fallback value that one-shot routing cannot see, the Nexus becomes a feedback loop rather than a classifier.

What remains unsolved

The residuals are specific.

The first obligation is the full task set with DRIFT measured. Twelve tasks and DRIFT remain pending, and every conclusion above is bounded by their absence.

After that: the LLM router needs scored runs before any claim about generative routing, and live sequential episodes need running before the stop-policy analysis means anything.

Seeding control is designed but untested — whether the Nexus should own where activation enters the graph, given that Chapter 5 found branch separation happening at seeding.

Model routing, human escalation, and any reinforcement-learned router wait behind the simpler policies. Temporal validity, forgetting, and consolidation stay in their own chapters. The registry is built to admit them when they arrive.

A system with several ways to remember now has a policy for choosing between them, and the policy’s value so far is measured in cost and harm avoided rather than answers improved. Whether a control layer earns more than that remains unresolved on this evidence.

References