Learning From Experience
Remembering the past and learning from outcomes are different operations; adaptation earns itself only when credit can be assigned and replay survives.
Chapter 15 asked how a long-lived memory system keeps accumulated history manageable without destroying what later work may still need. That remains memory management: the system changes which retained experiences are represented, available, or allowed to compete.
This chapter crosses a different boundary.
Suppose the system performs a migration, observes that it succeeded, and then changes what it will retrieve next time. Suppose it notices that the same sequence of checked actions has worked repeatedly and turns that sequence into a reusable method. The system is no longer only asking what the past contains or which part matters now.
It is allowing the outcome of present behaviour to change future behaviour.
That is learning from experience.
The distinction matters because a memory system can be excellent without being self-modifying, and a self-modifying system can become worse precisely because it mistakes correlation for evidence. Outcome adaptation and procedural memory belong together because both ask the same deeper question:
When is experience strong enough to justify changing how future work will be done?
Memory does not require that return edge. Learning does: an evaluated outcome is allowed to justify a change in how later situations will be processed. That closes the loop:
flowchart LR
M[Memory] --> A[Action]
A --> O[Outcome]
O --> AT[Attribution]
AT --> U[Policy or procedure update]
U -.->|"the edge this chapter must earn"| M
linkStyle 4 stroke:#3978c5,stroke-width:2px,color:#3978c5
The forward path shows memory influencing action and producing an outcome. Attribution then asks what, if anything, deserves credit. The dashed return edge is the candidate learning step, and the rest of the chapter asks when that step is justified.
Memory is not yet learning
The book’s behavioural definition of memory is:
retained past
↓
changes present behaviour
Learning adds a return transition:
No neural weight update is required for that distinction. The change may live entirely outside the model: a retrieval policy, an availability rule, a promoted procedure, or another versioned decision rule. Merely recording a correction or exception is not enough; learning requires evaluated experience to change how later situations will be processed.
But the causal structure changes.
Memory asks whether the past influenced the present.
Learning asks whether the consequences of the present should alter the mechanism that influences the future.
The second question is harder because success and failure do not explain themselves.
Outcome observed is not outcome attributed
Imagine a migration succeeds after the system retrieved five memories:
- the PostgreSQL decision;
- the compatibility-first warning;
- an old incident;
- a general migration note;
- an irrelevant runbook echo.
Which memory deserves credit?
The tempting rule is usage:
memory retrieved
↓
task succeeds
↓
raise memory priority
That rule creates a feedback loop:
retrieved because highly ranked
↓
present during success
↓
strengthened
↓
retrieved even more often
A lucky early success can lock in a mediocre memory. An irrelevant memory can ride beside the real cause and accumulate credit. A stale memory can survive because a capable reader worked around it. A useful warning may be ignored and receive no credit even though it should have changed the plan.
The attribution distinctions make those failures precise:
Each distinction does real work:
Retrieved means the memory entered context.
Used means the observable output drew on it.
Followed means the action complied with it.
Relevant means it bore on the task whether or not the agent followed it.
Causally helpful means the outcome would have changed without it.
Correlated with success means it was often present when successful outcomes occurred.
Of these signals, intervention-backed evidence of causal helpfulness offers the strongest basis for adaptation. Retrieval, use, or co-occurrence with success does not by itself assign credit.
This chapter therefore refuses the fiction that a scalar reward can be attached cleanly to every memory after every task.
Chapter 12 supplied the instrument
Chapter 12’s matched-intervention design — no memory, correct memory, remove decisive memory, restore decisive memory, wrong or stale memory — produces intervention evidence about whether supplied memory affected scored behaviour. On the six-task remove/restore subset, success moves from 0.250 under removal to 0.778 under restoration. Those paired conditions support attribution to decisive memory on the constructed tasks; the 0.524 oracle and 0.042 wrong-memory means come from different condition subsets and should not be treated as one directly matched ladder.
So the conservative posture in this chapter is a judgement about cost and scope rather than the complete absence of an instrument. Chapter 12 shows that intervention evidence can be obtained when decisive items and controls are known in advance, but its verdict remains bounded to a small controlled task set. It does not establish practical attribution for ordinary work, where several memories may contribute partially and no single removal is decisive.
Any adaptation policy should therefore use intervention-backed behavioural attribution evidence where available and must not manufacture causal credit from ordinary co-occurrence. Where such evidence is absent, utility updates remain hypotheses rather than established credit assignments.
What outcomes are allowed to change
Chapter 15’s separation between truth-confidence and utility now has direct consequences.
An outcome may legitimately alter:
expected utility
retrieval priority
availability preference
procedure candidacy
exception set
policy proposal
An outcome does not automatically alter:
truth
historical fact
provenance
supersession
what actually happened
A memory that worked three times may be more useful. It is not therefore more true.
A true memory that was irrelevant to ten recent tasks does not become false.
This keeps confidence and utility separate:
truth-confidence
changed by evidence
expected utility
changed by task/outcome evidence
Without that separation, a self-improving retrieval system can rewrite epistemology as popularity.
Backtest before promotion
Chapter 10 produced one of the book’s strongest control results. A policy change that looked locally correct scored −0.074 on the primary metric when replayed over the broader historical task set. It was rejected.
Chapter 11 repeated the pattern exactly. Dropping the cancelling-evidence leg would have recovered a candidate the gate was suppressing — locally correct, again.
The replay scored −0.354 and breached two gates. The harmful-task rate went from 0.000 to 0.500. Correct abstention went from 1.000 to 0.000.
Twice now, the repair suggested by the failure in front of the system has been the wrong one.
Those replay failures motivate a conservative candidate protocol — outcome observed, candidate change, historical replay against the primary metric plus regression gates, then promote or reject — rather than the direct alternative:
outcome observed
↓
change live policy
Replay matters because an adaptation proposed from one recent outcome can overfit that experience and regress on earlier task classes.
A policy update may repair the task that inspired it and damage an earlier class. A retrieval preference learned from architecture work may crowd out publication evidence. A rule learned from three successful migrations may fail catastrophically on a fourth because one precondition differs.
Earlier chapters introduced versioned ContextPolicy, gate policy, and assembly policy, and tested policy revisions through replay. Any outcome-driven adaptation would need the same discipline.
Every promoted adaptation needs:
proposal
evidence
policy version
evaluation set
regression gates
promotion decision
rollback path
Under that design, learning would become inspectable change management rather than invisible reinforcement.
The exploration problem
There is another asymmetry.
A memory that is never retrieved cannot receive positive outcome evidence. A policy that already prefers A over B will collect more experience with A, which can make A look increasingly justified simply because B is no longer tried.
This is the classic shape of an exploration problem, but the book does not need to build a reinforcement-learning system to acknowledge it.
It does need to reject a misleading inference:
“Never used” is not the same as “not useful.”
A static or conservative policy may sometimes be preferable to an adaptive one because it avoids locking the system into early accidents.
That possibility must remain a valid verdict.
From useful memory to reusable method
Outcome adaptation changes which memories are likely to influence later work.
Procedure extraction goes further. It asks whether successful experience can be represented as a reusable method.
Compare:
Compatibility mode was important to the migration.
with:
For this class of migration:
1. enable compatibility mode
2. apply the schema change
3. shadow-read both versions
4. validate
5. remove compatibility only after checks pass
The first is a declarative memory. It can inform a future decision.
The second can organise action.
That difference is why procedure is the strongest form of the book’s behavioural claim — and also the most dangerous.
A stale fact can mislead. A stale procedure can perform the wrong action.
Procedures are derived action claims
A candidate procedure representation can stay small:
goal
preconditions
steps
checks
evidence
Goal says what kind of work the procedure serves.
Preconditions define the applicability envelope.
Steps preserve the sequence.
Checks prevent blind replay by gating progress on observed state.
Evidence preserves the episodes, documentation, postmortems, or experiments from which the procedure was derived.
The candidate representation stays deliberately small. Branching languages, composition operators, parameter systems, autonomous skill graphs, and workflow DSLs remain outside this proposal until a measured failure requires them.
A procedure is best treated as a derived claim about action.
That means it inherits all the discipline of other derived state:
- provenance;
- scope;
- validity;
- versioning;
- supersession;
- re-evaluation when evidence changes.
Preconditions bound applicability
In this candidate representation, the most important field is not the step list.
It is the precondition.
A remembered procedure may have worked under one platform version, one service topology, one data constraint, or one compatibility contract. Similar future work can differ on exactly the hidden condition that made the old procedure safe.
Blind replay is therefore a failure mode this design must guard against.
The Strategy X incident (incident-034) makes the failure concrete. The procedure disabled foreign-key checks before copying rows, then was meant to restore them. An exception path skipped that restoration, leaving the checks disabled and permitting orphaned rows. The missing safety condition was not merely “run this on a similar migration”: Strategy X required an exception-safe restoration guard and a clean foreign-key validation before it could be considered safe to run.
Remembering the successful-looking sequence without that prerequisite would preserve the steps and lose the reason the procedure failed. A future system must surface the failure and its checks together, not convert repeated shape into permission to replay.
Under this candidate design, a procedure is available for action only when its relevant preconditions can be checked.
That introduces a natural boundary: the memory system may remember a procedure even when the agent lacks the tools or world-state access to verify its preconditions. In that case the correct behaviour is not confident replay. It is to request evidence, defer, or use the procedure as advisory context.
Memory should not silently convert a remembered method into authority to act.
Checks distinguish procedure from script
The second safety feature is validation during execution.
A script says:
do A
do B
do C
A robust remembered procedure says:
do A
check X
if X holds:
continue
otherwise:
stop / roll back / seek evidence
The book does not need a general branching language to preserve this distinction. A linear sequence with explicit validation gates is enough for the first experiment.
That restraint matters. The moment procedures become arbitrary programs, the chapter has become a workflow-engine design project rather than a memory investigation.
A combined experiment
This chapter has no dedicated outcome or procedure run. The design below is an untested boundary proposal, and it asks one broader question:
Does feeding evaluated experience back into memory improve the next task, and if so, is the useful change a retrieval preference, a reusable procedure, or neither?
A controlled repeated-task world would separate a genuinely necessary memory, a correlated-but-irrelevant fellow traveller, an actively misleading memory, a conditionally useful memory, a successful action sequence with explicit preconditions, a near-miss task where the same sequence becomes unsafe, early lucky outcomes favouring the wrong memory, and a later environmental change invalidating the old procedure. STATIC, USAGE, OUTCOME, INTERVENTION-BACKED, PROCEDURE-CANDIDATE, and PROCEDURE-GATED conditions would compete on that world, measured separately on task success, harmful-action rate, lock-in, recovery after contradiction, false strengthening, procedure transfer, near-miss refusal, rollback, and evaluation cost.
Book hypothesis. Outcome feedback can improve future memory selection only when attribution is strong enough to distinguish useful evidence from fellow travellers, and procedure extraction can improve repeated task performance only when applicability conditions and checks survive transfer. If those conditions cannot be established reliably, static traceable memory should remain the default and automatic adaptation should be deferred.
That is a much weaker claim than “memory should learn from every outcome,” and deliberately so.
What would count as a negative result?
Several negative outcomes would be useful.
If ordinary static memory plus a strong reader matches every adaptive condition, adaptation has not earned its statefulness.
If usage reinforcement locks onto early noise, it should remain a baseline pathology.
If intervention-backed utility updates help selection but procedure extraction adds no behavioural gain, only the first half survives.
If procedures improve mean success but cause rare destructive replay, they fail even if an aggregate rises.
If only human-authored procedures remain safe, automatic extraction should be demoted to candidate suggestion.
If outcome evidence cannot support causal attribution at practical cost, the chapter should end with a boundary rather than a mechanism.
That boundary would be valuable:
Memory can preserve outcomes and expose them to evaluation without automatically rewriting itself from them.
Where memory ends
The distinction can now be stated cleanly.
Memory is durable causal dependence on a retained past.
Learning is modification of the mechanism that determines how future situations will be processed.
A system may have rich memory and almost no autonomous learning. It may keep project history, resolve current truth, preserve open work, choose relevant context, and reuse human-authored procedures without updating any policy from its own outcomes.
That is a coherent architecture.
This chapter therefore treats automatic adaptation as optional. It should enter the final system only if an experiment demonstrates value that static, auditable memory cannot obtain more simply.
Research foundations
Several agent systems demonstrate pieces of this possibility.
Reflexion carries verbal feedback across attempts. ExpeL derives transferable lessons from successes and failures. Voyager maintains an expanding skill library and uses environmental feedback to improve later behaviour. Self-Refine applies generated feedback iteratively.
These systems show that prior outcomes can change later performance. They do not eliminate the attribution, lock-in, applicability, and provenance problems this chapter makes explicit.
Any future adaptation would still need to be versioned, replayable, reversible, and separate from truth maintenance.
References
- Reflexion: Language Agents with Verbal Reinforcement Learning (2023).
- ExpeL: LLM Agents Are Experiential Learners (2023).
- Voyager: An Open-Ended Embodied Agent with Large Language Models (2023).
- Self-Refine: Iterative Refinement with Self-Feedback (2023).
The handoff
That restraint is the line the next question needs. Once retained history can change behaviour — and the behavioural instrument has shown that it can, in both directions — memory becomes an authority channel, and authority must be priced before any architecture is declared final.