Let Old Context Fade
Explore progressive fidelity and tiered representations instead of cliff-edge truncation or one-shot compaction.
A long session crosses its compaction threshold mid-investigation. The agent is halfway through a dependency migration: three candidate approaches tried, one rejected for a subtle version conflict, the current attempt half-applied, the exact error that killed the second attempt still visible three turns back. The compactor fires, because tokens crossed a number, and replaces 40,000 tokens of working history with a 2,000-token summary. The summary is fluent and mostly accurate. It records that an approach was rejected without recording which version conflict killed it. Six turns later the agent, unable to see the conflict, retries the rejected approach. The threshold did its job. The session did not survive it.
Chapter 11 built the machinery that produced this failure and measured why summaries fail this way. This chapter asks whether the underlying shape is necessary. Must information jump in one step from full fidelity to a single compressed representation, or can the same episode exist at several explicit levels of detail, with rules governing which level is active? The question matters because the cliff confounds two decisions that deserve separation: what may be lost, and how much of what remains is shown right now.
One content, several representations
One distinction has to come before anything about tiers, because the rest of the book leans on it. The information an episode carries and the representation in which it is shown are different things.
The incident report a session investigated is one piece of information. It may exist as the full text, as a dense rewrite, as a compact summary, as an anchor of a few dozen tokens, or as a bare reference. Those are five representations of one content, and choosing among them is a decision separate from deciding whether the incident matters at all.
Once that is stated, earlier chapters read differently. Chapter 11 argued that some content must survive exactly, which is a statement about which representations of it are acceptable. Chapter 13 will separate an anchor from a pointer, which are two more representations. Chapter 18 generalises the point to prose, records and tables.
For a mechanism that has to assemble a bundle the split has a practical consequence. It must know that the full text and the compact form are alternatives to one another and not two items that happen to overlap, or it will spend budget saying the same thing twice.
The other half is a bound. Each content has a lowest representation it is allowed to take, below which what remains is no longer the same information for the purposes of the task. The safety rule’s lowest is its exact wording. A decision’s is its outcome and its reason. A transient log’s may be nothing at all. That lowest permitted form is the floor. It is what pressure pushes against, and what stops the push.
Fidelity is not retention
Chapter 7 classified what transformations each item permits. Chapter 12 introduces a second variable that the classification alone does not determine:
RETENTION CLASS
What transformations are allowed?
FIDELITY LEVEL
Which allowed representation is shown now?
A decision may be COMPRESSIBLE, which permits semantic transformation, while currently rendered at full detail because pressure is low. The same decision under severe pressure may render as outcome plus rationale, then as outcome alone, without any change to its retention class. Conversely, a PIN item never descends at all: its retention contract fixes its minimum fidelity at exact, and no schedule overrides that. The relationship is one-directional and worth stating as law:
A fidelity policy may select among representations the retention contract permits. It may never select one the contract forbids.
Pressure demotes. Floors detain. Confusing the two, treating a demoted item as a reclassified one, destroys the audit trail Chapter 10 built: the record must show that an item rendered compactly is still COMPRESSIBLE, not that it became DISCARDABLE.
That relationship gives the floor its place. A fidelity floor is the lowest representation an item is permitted to reach: exact and verbatim for a constraint, decision plus rationale for a choice, outcome plus reason for a resolved investigation, possibly very small for transient operational detail. Floors derive from retention semantics and probe results, never from intuition. The compiler of Chapter 23 takes one floor per candidate as an input and enforces it. Where the floor comes from is the builder’s problem, and no repository derives floors yet.
Age is not permission
Chapter 10 established that old does not mean discardable. The parallel statement holds here with equal force: old does not mean low fidelity. A six-month-old architecture constraint may require exact representation while a five-minute-old compiler log tolerates aggressive reduction. Any policy shaped like older-automatically-shorter repeats the age fallacy at a new level of sophistication, demoting the governing past to subsidise the chattering present. Therefore:
Age may influence demotion only after retention semantics establish that demotion is legal.
Age enters the machinery below as a priority input among others, never as a licence. What age cannot do, pressure and relevance sometimes can, which the next sections separate carefully.
Three questions, one pipeline
Progressive fidelity decomposes into three independent decisions that the chapter refuses to merge:
Eligibility. May this item lose detail at all? Chapter 7 answers most of it through retention class and floors.
Level. Which representation fits current conditions? A choice among legal tiers under pressure, relevance, and cost.
Timing. When is the representation change applied? Chapter 9’s economics govern here: a demotion eligible now may be cheapest applied later, at a cache-safe boundary, on prefix expiry, or at a milestone transition.
In pipeline form:
item
↓
legal fidelity range
↓
desired tier
↓
safe application boundary
↓
rendered representation
No compiler is built from this pipeline. It is a requirements sketch showing where each earlier chapter plugs in: retention supplies legality, budget and relevance supply desire, cache state supplies timing. A design that fuses the three into one score repeats Chapter 7’s importance-score error with more variables.
Generate once, render many times
The most consequential architectural pattern in this chapter separates two events that naive systems fuse: creating reduced representations, and choosing which one to show. Magic Context’s current architecture, verified against its primary source, implements exactly this split. Its historian emits each history compartment with four detail tiers at generation time, plus an importance score, facts, and events. Later, a deterministic decay renderer selects one tier per compartment per pass with no model call, with the tier chosen from age, importance, and budget pressure. Raw messages persist separately for exact expansion. Every mutation replays byte-identically across passes for cache stability, and re-tiering happens only on defined fold boundaries rather than continuously.
Three properties of that design deserve isolation from its product specifics. First, demotion is not itself a generative event: choosing anchor-only for a compartment calls no model and invents nothing, which removes an entire class of variance that per-turn re-summarisation would introduce. Second, the tiers are generated once from source rather than cascaded from each other, so lower tiers do not inherit the errors of higher ones; whether source-grounded generation measurably beats cascaded generation is one of the experiment’s axes, not an assumed result. Third, the raw range survives for expansion, which makes every demotion reversible in a way Chapter 11’s summaries, standing alone, are not. The book adopts none of Magic Context’s four tiers as canonical, its importance score stays one implementation’s control variable per Chapter 7’s refusal, and its memory subsystem stays out of this chapter entirely. What transfers is the split: generative loss events at creation time, deterministic render events afterwards.
That split is also what separates this chapter from repeated compaction. Chapter 11 studied chains where each summary compresses its predecessor and loss compounds by construction. A tier system that regenerates every tier every turn has simply hidden that chain inside a scheduler. The firewall reads:
regenerate-then-select-every-turn
=
repeated compaction wearing tiers as costume
Genuine progressive fidelity creates representations on a slower cadence than it selects among them, and the experiment below is designed to catch systems that fake the distinction.
What tiers contain
Tiers defined by token count alone repeat the ratio fallacy of Chapter 11 at finer granularity. A shorter tier can preserve constraints, decisions, identifiers, status, and provenance while dropping narrative; a longer tier can still omit the decisive fact. Token budget constrains a level. Information survival defines it. The chapter therefore specifies tiers by survival contract rather than size, derived from Chapter 7’s dimensions. Narrative detail may decay early. Exact identifiers, decision rationale, failure history, locational anchors, provenance, and uncertainty status each carry floors set by retention semantics and probe results: a file line number may vanish before the subsystem name, but the reason the chosen approach won outlives the incidental failed command that tested it. These orderings are hypotheses for fixtures, not laws; the experiment section makes them falsifiable per class.
Two structural consequences follow. First, not every item uses every tier. Some items admit only exact representation. Some admit full and compact but nothing between. Forcing all items through an identical four-rung ladder manufactures transitions nobody needs; retention semantics precede decay, and the tier set per item is whatever its legal range supports. Second, tiers need not be prose at decreasing lengths. A structured canonical record rendered deterministically into detailed, compact, and anchor views can protect exact fields and provenance by construction, at the price of an intermediate representation that may itself be wrong and of narrative relationships that fields capture poorly. Independent prose tiers preserve narrative at the price of cross-tier contradiction risk. The experiment compares both without assuming a winner.
The decay dimensions deserve one level more concreteness, because floors are set per dimension rather than per tier. Narrative detail, the blow-by-blow of how an investigation unfolded, decays first almost everywhere: once an outcome and its reason survive, the intermediate steps are usually safe to shed. Exact identifiers, file paths, revision markers, configuration values, decay last or never, keyed to exactness rather than age. Decision rationale outlives the search trace that produced it; failure history outlives the passing curiosity that read it, because rejected alternatives constrain future choices. Locational anchors, subsystem names and stable references, persist after line numbers fade: a line number locates text that edits invalidate, while the subsystem name locates responsibility that persists. Provenance and uncertainty status ride at whatever tier carries their fact, never below it. None of these orderings is law. Each is a probe hypothesis: the fixtures test per-dimension floors by demoting dimensions independently and watching which demotions break hidden tasks. Where a floor holds across traces, it graduates into policy. Where it varies, the variance itself is the finding, and the tier contract stays conservative.
Anchors deserve precise definition since the term is easy to use loosely. An anchor is not a summary of an episode. It is the minimum representation preserving enough identity to recognise what happened, why it may matter, and which larger episode it belongs to: a stable episode identifier, a short outcome, the critical invariant, canonical entity names, a source identifier. An anchor must never pretend to contain detail it lacks. Its honesty is what makes it safe as a lowest resident tier rather than a decorative stub.
The tier fan for one episode looks like this, with the caveat that most items use a subset, never the full fan, of these levels:
SOURCE EPISODE
│
┌───────────┼───────────┐
↓ ↓ ↓
FULL DENSE COMPACT
│
↓
ANCHOR
retention policy defines legal floor
budget and relevance select active view
Pressure creates necessity; age merely prioritises
A quiet session at 20K of a 1M-token window has little reason to demote anything for age alone. A session at 190K of a 200K window may need strong reductions regardless of recency. Hence the chapter’s operating rule, stated as testable rather than settled:
Age can influence priority; budget pressure creates necessity.
Current relevance complicates both, modestly and under control. An old episode can become central again; a recent one can already be complete. The fixtures manipulate known relevance directly, old-but-needed versus old-and-done versus recent-complete versus recent-active, and watch whether age-only and retention-aware schedules diverge. No retrieval or ranking machinery is built from this; relevance here is a controlled experimental input, and discovering relevance in production belongs to later chapters. Promotion back up the tiers is allowed conceptually, an old anchor re-rendered densely when its episode reactivates, with the discovery mechanism explicitly deferred. Demotion must never be modelled as semantically permanent, which is one more reason lineage and source identity travel with every tier.
Policy first, schedule second
A final separation keeps the machinery honest. Fidelity policy declares which representations are legal; decay schedule decides when the active representation changes. Policy comes first and schedule cannot override it. A PIN item’s legal set is exact-only under every schedule and every pressure reading. A COMPRESSIBLE episode’s legal set might span full, dense, compact, and anchor, with age, relevance, and budget pressure selecting among them turn by turn. Confusing the two produces characteristic errors in both directions: a schedule that demotes a PIN item because pressure is high has mistaken urgency for permission, while a policy that refuses all demotion because importance is high has mistaken value for presence, Chapter 7’s error in new clothes.
The separation also clarifies what each part of the experiment tests. Tier-legality questions, does this item admit compact representation at all, are answered by the first test’s static survival measurements. Schedule questions, when should the legal demotion actually render, are answered by the second test’s trajectory runs under matched budgets. A policy can be correct while its schedule is wasteful, demoting too eagerly and paying cache churn for no behavioural gain, or too timidly, carrying full detail through pressure that justified compact views. Scoring the two separately is what lets a negative result land precisely: keep the tiers, fix the schedule, or the reverse, instead of discarding the whole mechanism over one half’s failure.
Deterministic rendering, hysteresis, and the cost of switching
Once tiers exist, choosing among them should be reproducible under identical state. Deterministic rendering buys repeatability, auditability, stable experiments, lower management cost, and freedom from per-turn generative drift. It does not buy correctness: a deterministic rule can reliably select the wrong tier, and the chapter measures decision reproducibility separately from decision quality throughout. Both need numbers, because a policy debate conducted only in quality terms cannot distinguish a bad rule from a noisy one.
Switching itself has a price, and the chapter prices it twice. First in cache economics, inherited whole from Chapter 9: every tier change is a potential prefix mutation with a divergence position and a radius, so desired demotion and applied demotion stay separate, with deferral to cache-safe boundaries, prefix expiry, or milestones. A renderer that re-renders history differently each turn because age ticked one minute creates continual mutation for no behavioural gain; discrete tiers switched at boundaries dominate continuous micro-rewriting on cost alone, before fidelity is even considered. Second in stability: pressure fluctuating around a threshold can thrash tiers downward and upward alternately, so a minimal hysteresis rule, separate demotion and promotion thresholds so small fluctuations do not rewrite the bundle, earns its place as runtime plumbing rather than framework. Small, plainly defined, measured for thrash reduction rather than admired.
Tier changes also bill their creation. Generating several representations costs more upfront than generating one summary: historian input, tier output, storage, decision computation, rendering overhead, cache impact, all counted against downstream savings. A four-tier mechanism is not free because its live rendering is compact. The experiment’s overhead accounting from Chapter 11 carries over unchanged.
How many tiers earn their place
The chapter refuses to begin from any plugin’s tier count. Two, three, four, or more levels are hypotheses, and the experiment is allowed to collapse them. The survival criterion is Pareto-shaped: a tier survives only if it creates a useful fidelity-versus-cost operating point that neighbours do not. A tier costing 600 tokens that performs no better than a 350-token neighbour is dominated and should disappear. A tier saving 50 tokens while destroying survival has priced itself out. Every tier must answer two questions, what it preserves that the next-lower tier does not, and what cost it saves relative to the next-higher tier, or be merged away. This is also where excess fidelity gets measured: tokens spent above the minimum sufficient oracle tier, punishing timid never-demote policies symmetrically with aggressive ones. Under-fidelity selects below the sufficient tier; over-fidelity selects above the necessary one; the oracle defines both only where future probes make sufficiency checkable, never as production labels.
Functional layering enters here as a deliberate non-conflation. Recent hierarchical work separates planning from execution detail into isolated layers with their own summaries, which is organisation by function, not gradation by fidelity. A planning layer may itself need high or low fidelity; an execution layer likewise. The chapter records the distinction because both ideas will coexist in later systems, and a planning tier mistaken for a fidelity tier inherits the wrong contract. Similarly, learned external managers show that the preferred operating point can depend on the consuming model: stronger agents exploiting longer raw contexts, weaker ones needing aggressive distillation for reliability. That capability-indexed trade-off is reported as a preprint finding that informs the experiment, through a same-schedule-across-two-models extension kept strictly optional, rather than as architecture the book adopts. No importance score sneaks back in through either door.
Proposed experiments
The question. Do explicit tiers, generated once and selected deterministically, beat a single compaction at matched budgets, and do they earn each tier?
The design, in brief. Use deterministic episodes with exact constraints, identifiers, decisions with rationales, rejected alternatives, unresolved work, uncertain hypotheses, provenance and incidental detail, plus hidden probes the scheduler never sees. First test the tiers themselves: raw source, a flat summary, tiers generated from one another, tiers generated independently from the source, a structured record rendered deterministically, and an oracle, at matched token budgets, scoring exact, semantic, status and provenance survival and cross-tier contradictions. Then test schedules over long trajectories under matched active budgets: full history, hard truncation, flat compaction, tiers scheduled by age alone, tiers scheduled by retention and budget, and an oracle schedule. Score the policy and the schedule separately, so that a negative result lands on the half that failed.
The measurement that matters. Survival per token, by information class, never averaged across classes; and decision reproducibility reported apart from decision quality. A tier earns its place only if it creates an operating point its neighbours do not.
What would change the book. If a flat compaction matches tiered rendering at equal budgets, or if only two levels earn their place, the tier machinery is removed and the chapter narrows to that. Flat summaries that beat tiers, tiers that cost more to create than they save, or frequent contradictions between tiers would each remove the motive. Nothing here has been run.
What a builder hands the compiler
Nothing in this chapter is a component. What it specifies is what a system that builds candidates must produce for the compiler of Chapter 23 to use it.
For each content, the representations it may be shown in, each with its token count and its rank from bare reference up to full text, and the lowest rank the content is allowed: its floor. The tiers are generated once, from the source, at a slow cadence. Which one is shown for a given computation is chosen later, deterministically, by the compiler, from the forms it is given and under the floor. The compiler does not write a compact form when only a full one exists, and it does not shorten one to fit.
That is the reason the generate-once, render-many split is right. Loss events that generate text belong at creation, where they can be measured once. Selection should not create them.
The fixtures that would test any of this need episode boundaries, several variants per episode, hidden probes and progressions of budget pressure. Real long traces would then answer the ecological questions fixtures cannot: how much old context exists, which classes dominate it, and how often each tier would be chosen.
Residency is the boundary
Two firewalls close the chapter. First, an anchor with semantic content resident in the window is Chapter 12; a pointer whose useful content lives elsewhere is Chapter 13. Low-fidelity resident does not equal externalised, and tier-to-zero is pruning owned by Chapter 10, never a fidelity level. Second, fading live history is not memory: representation inside current computations stays here while durable cross-session influence stays in the Memory book, with Magic Context’s memory subsystem cited only as a boundary. Freshness stays in Chapter 20, authority in Chapter 19: an old tier is not a false tier, and a compact rule outranks a verbose log regardless of size.
What remains in the window after tiers have done their honest work is the irreducible remainder: pinned exact spans, active dense episodes, compact views of the recent past, anchors of the deep past. Still resident, still costing capacity and cache stability. The next move is not another representation. It is departure:
Progressive fidelity still consumes window space. The next move is to take information out of the window entirely.
References
- cortexkit. “Magic Context.” Context plugin, MIT at review, verified September 2026. Four detail tiers per compartment; deterministic decay rendering; raw retention for expansion. https://github.com/cortexkit/magic-context
- cortexkit. “Magic Context ARCHITECTURE.md.” Primary source, verified September 2026. Generation/render split; replay determinism; fold boundaries. https://github.com/cortexkit/magic-context/blob/master/ARCHITECTURE.md
- Wang, X., Xiao, J., Cui, S., et al. “HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation.” Preprint, arXiv:2608.15703, August 2026. Functional layers vs fidelity gradation. https://arxiv.org/abs/2608.15703
- Yi, L., Lei, R., Yao, L., et al. “Learning Agent-Compatible Context Management for Long-Horizon Tasks.” Preprint, arXiv:2605.30785, May 2026. Fidelity-reliability trade-off; capability-indexed operating points. https://arxiv.org/abs/2605.30785