← DSPy From First Principles

Search, Memory and Long Context

Hold the task, rewrite program, model configuration and metric fixed, vary how supplemental evidence is acquired, and measure six context policies — including adaptive selection and a failed RLM path.

Chapter 15 measured an agent that had every advantage. Two files, one defect, a bounded tool surface, a budget it used well. It found both implicated files, held both halves of the bug in context, and still stopped one step short. That was a reasoning failure, and Chapter 17 exists to attack it.

This chapter is about the problem that sits before the reasoning problem. The fixture had two files. A real repository has thousands, and a bounded tool surface does not tell the program which of them to read. Something has to decide what evidence the program even looks at, and that decision is not one mechanism. It is at least five, and they are routinely collapsed into a single word.

1  put the evidence directly in the prompt
2  retrieve evidence before the model runs
3  let the program retrieve its own evidence while it runs
4  retrieve past cases instead of current files
5  let the model explore a large structure programmatically

Each solves a different failure, each has a different cost, and — this is the chapter’s method — each is an experimental variable that can be held against a fixed program and a fixed metric so that a score difference means something.


1. One target, one program, one metric

The temptation with a comparison like this is to run six policies on the full corpus and read a leaderboard. That produces a number per policy and tells you almost nothing, because the policies do not differ only in quality — they differ in what information they are allowed to touch, and a leaderboard hides that.

The experimental factor is the evidence-acquisition policy. The final rewrite program and its task-facing inputs stay fixed, while the policies differ in how they select supplemental evidence and in the cost required to select it:

  • One target case. ed-003, a canonical development case. Anna is warning her brother; the sentence I do not think we should go there, Anna said, because it is unsafe. needs dialogue punctuation. It is walked through once, not swept across the corpus.
  • One rewrite program. dspy.Predict(RewriteWithEvidence) — the ordinary task inputs plus one advisory evidence field that may be empty. Every policy eventually calls this same program.
  • One scoring setup. The frozen editorial_metric_v1 from Chapter 8 supplies the primary score; v2 is recorded alongside it as a secondary semantic check. It is not independent evidence in the strong sense because the semantic judge uses the book’s configured model family.
  • One task-model configuration. The canonical local qwen3:latest at temperature 0 executes the final rewrite. Adaptive selectors use separate dspy.LM objects with the same recorded configuration.
  • The holdout is sealed. The seven holdout cases remain outside the run; ed-004 is one of them.

2. Evidence role is not evidence quality

Before the numbers, the distinction that makes them readable.

The six rows do not all have the same evidence role. Direct context adds no supplemental packet. Four policies are intended to add decision-time evidence. labeled_training_memory deliberately exposes a richer source: labels from other training cases.

PolicyWhat it feeds the programRole
Direct contextno supplemental evidence; the case’s context field is already in the signaturecurrent case
Fixed guideline retrievaltop-2 lexical hits from a small editorial-guideline corpusdecision-time guidelines
Decision-time case memorytraining-case inputs selected without their reference rewrites; in this measured run it selected no supplemental casedecision-time prior cases
Labeled training memorytraining cases including their reference rewriteslabeled training data
Agent guideline retrievalguidelines the program selected for itself via toolsdecision-time guidelines
RLM guideline explorationguidelines selected by programmatic corpus explorationdecision-time guidelines

labeled_training_memory is the row to keep separate. The reference rewrites for ed-001 and ed-002 are not the answer to ed-003, so merely showing old labeled examples is not target-answer leakage by definition. Such examples can be legitimate demonstrations in a system whose deployment policy permits them.

They are nevertheless a different information privilege from the decision-time policies in this experiment. The runner therefore records evidence_role="labeled_training_only", and the comparison must not attribute any advantage from that richer role to the retrieval mechanism alone.

A gain produced under a richer evidence role is not evidence that the retrieval policy itself is better. If the declared deployment protocol forbids that role at inference time, the row is inadmissible for a deployment-equivalent comparison.

Chapter 18 turns that distinction into enforcement: the runtime may consume only evidence roles and sources declared by the experiment. This chapter keeps labeled training memory visible as a deliberately non-equivalent comparison arm.


3. The measured comparison

One LM-backed run, ed-003, six policies, the same rewrite program and metric throughout. Token counts are from the run manifest; v1 and v2 agreed on every row.

PolicyEvidence charsSelection tokensTotal tokensElapsedScore (v1 = v2)
Direct context004172.4 s0.8769
Fixed guideline retrieval33704742.1 s0.8769
Decision-time case memory004081.9 s0.8769
Labeled training memory13804672.4 s0.8769
Agent guideline retrieval2043,8204,28419.7 s0.6500
RLM guideline exploration004082.0 s0.8769*

* The RLM selection step crashed before selecting anything; this row is the rewrite program running with no supplemental evidence, not an RLM result. Section 7.

The table already rules out one tempting reading. The four 0.8769 rows did not feed four distinct supplemental-evidence strings: direct context and decision-time case memory both reached the rewrite with an empty evidence field, while fixed retrieval and labeled memory supplied different packets. The failed RLM path also reached the rewrite with an empty evidence field.

The adaptive agent row did incur a clear cost. It used 3,820 selector tokens and 4,284 total tokens, versus 474 total tokens for fixed retrieval — about 9.0× as many total tokens. Its 19.7 seconds versus 2.1 seconds is about 9.4× the elapsed time in this run.

The important question is therefore not whether the agent “lost” by 0.2269. It is why the frozen metric assigned 0.6500 to the particular rewrite it produced.


4. The tie is the metric, not the evidence

Start with the four-way tie, because it is easy to misread as “evidence selection does nothing here.”

ed-003 does have numerical headroom: 0.8769 is below the metric’s theoretical maximum. The four-way tie therefore cannot be explained by saying the score had nowhere to move. What happened is narrower: several different output strings mapped to the same v1 component values.

There is also a built-in negative control hiding in the table. Direct context and decision-time case memory both supplied the final rewrite program with the same task inputs and an empty supplemental evidence string, yet they produced different output strings. The case-memory selector returned no evidence in this run, so that output difference cannot have been caused by case memory. It is another instance of the execution-order or runtime-state variation Chapter 6 already warned about.

The rewrites still expose a real metric defect:

PolicyRewriteAdded quotation marks?
Direct / fixed / labeled"I don't think we should go there," Anna said, "because it's unsafe."yes
Decision-time memory / RLM fallbackI don't think we should go there, Anna said, because it's unsafe.no
Agent"I do not think we should go there," Anna said, "because it is unsafe."yes
Reference"I do not think we should go there," Anna said. "It is unsafe."yes

ed-003’s editorial goal is dialogue punctuation. The decision-time-memory call and the RLM fallback omitted quotation marks, while the other displayed rewrites added them. Yet v1 can still assign the same 0.8769 to punctuated and unpunctuated outputs because its changed component compares normalized lexical tokens and discards punctuation.

A punctuation-only edit therefore receives no changed credit, while contracting do not to don't or it is to it's changes the token sequence and earns that credit even if the punctuation goal is missed.

That deterministic scoring defect is the result to carry forward. It does not establish that case memory caused the unpunctuated rewrite — no case-memory evidence reached that rewrite call — and it does not establish that the tied retrieval policies were behaviorally equivalent.


5. A punctuation-only edit gets the 0.65 floor

The agent policy scored 0.6500. Decomposed against the actual metric components:

ComponentWeightResultContribution
changed0.35unchanged0.00
scope (length)0.25pass0.25
reference overlap0.401.000.40
Total1.000.65

The 1.00 overlap needs precise wording. It is perfect overlap under the metric’s normalized token-recall calculation, which ignores punctuation. The agent preserved the reference’s lexical tokens and added quotation marks, but it did not literally reproduce the human reference: the reference ends the speech tag with a period and begins a second quoted sentence, while the agent keeps one sentence joined by commas.

So this run does not establish that the agent produced the uniquely “most correct” rewrite. It establishes something cleaner and fully mechanical: a plausible punctuation-focused rewrite can receive the 0.65 floor because the metric’s changed component erases punctuation before comparing source and candidate.

That creates a perverse local incentive. A contraction such as do not → don't earns the full 0.35 change credit, while adding quotation marks without changing lexical tokens earns none. A candidate can therefore score higher for changing words than for performing the punctuation edit named by the goal.

The selector itself did find and read the relevant dialogue-punctuation guideline. Whether that selected evidence caused this particular conservative rewrite is not established by one trajectory, especially when the same task model is reused across sequential calls and Chapter 6 has already shown order-dependent variation.

The defensible finding is about the instrument: v1 can rank a punctuation-only edit below a lexical edit because punctuation is invisible to its changed component. The measured selector cost is a separate result.


6. Adaptation is not free, and here it bought nothing

The agent policy is dspy.ReAct over two tools — search_evidence(query) and read_evidence(id) — with a budget of four iterations. Its trajectory was short and sensible:

1. search_evidence("improve dialogue punctuation and rhythm, warning")
     → guide-dialogue-punctuation (overlap 4), guide-cause (1), guide-technical-scope (1), ...
2. read_evidence("guide-dialogue-punctuation")
     → the full guideline text
   selected: ["guide-dialogue-punctuation"]

It queried, ranked the dialogue-punctuation guideline first, read it, and selected it. The fixed lexical retriever selected that guideline plus guide-tone. On this five-item synthetic corpus, the agent’s narrower set is defensible, but one downstream rewrite does not establish that the selection was better.

The cost difference is measurable. The agent used 3,820 selection tokens and 4,284 total tokens, versus 474 total tokens for fixed retrieval: about 9.0× as many total tokens. Elapsed time was 19.7 seconds versus 2.1 seconds, about 9.4× in this run.

Adaptive model-driven retrieval therefore introduces a selection overhead that fixed lexical retrieval does not pay. That does not mean adaptive retrieval must always cost more end to end: on a larger task it may reduce downstream context, avoid unnecessary reads, or stop earlier than a fixed policy. The correct rule is to measure selector calls, evidence size, downstream calls, tokens, and latency together.

And “the program chose its own evidence” is still not the same as “the program chose better evidence.” Selection quality is measured behavior, not an assumed benefit.


7. Report the failed mechanism, not the fallback score

The RLM row scored 0.8769, and that number is a trap.

dspy.RLM is documented as experimental. It lets the LM explore external context programmatically by generating Python that is executed through a sandboxed interpreter, rather than placing the entire context directly in the prompt. DSPy’s default PythonInterpreter uses a Deno/Pyodide WASM sandbox, and this environment did not have a compatible Deno runtime:

CodeInterpreterError: Unable to determine the Deno version from 'deno'.
PythonInterpreter requires Deno >=2.0.0,<3.0.0.

The selection step raised before the model explored anything. The runner caught it, recorded selection_status: "failed" with the exception text, selected no evidence, and let the rewrite program run anyway — which is why the row scored the same as direct context. It was direct context.

The rule the run record enforces:

When a mechanism fails, the artifact says so. Crediting its fallback score to the mechanism is a lie the provenance is built to prevent.

0.8769 next to “RLM” would tell a future reader that RLM matched the baseline. What actually happened is that RLM never ran and the fallback matched the baseline because it was the baseline. The asterisk in the table is not a footnote; it is the result.


8. Every packet is fingerprinted, and that is what Chapter 18 needs

Every policy produces the same shape of record, whether it made zero model calls or crashed:

{
    "policy": "labeled_training_memory",
    "query": "...",
    "evidence_role": "labeled_training_only",   # the field that flags the leak
    "adaptive": False,
    "selection_status": "completed",            # or "failed", with the exception
    "selected_evidence_ids": ["ed-001", "ed-002"],
    "source_ids": ["chapter-opening-a", "chapter-opening-b"],
    "evidence_size_chars": 138,
    "selection_calls": 0,
    "selection_usage": {"total_tokens": 0, ...},
    "selection_trace": None,                    # or the full ReAct / RLM trace
    "packet_fingerprint": "…",
}

Twelve invariants are checked against those records and all passed — including historical_memory_uses_train_only, holdout_not_selected_as_case_memory, and holdout_not_inspected.

This is not bookkeeping for its own sake. Chapter 18 argues that you cannot audit an information boundary you did not record, and this chapter is where the recording starts to matter. Once a program can retrieve — guidelines, past cases, repository history, its own prior outputs — the question “did this result depend on something the program was not supposed to see?” becomes answerable only if every retrieved item carries provenance.

The labeled_training_memory row is the concrete example. Its evidence_role says labeled_training_only, so a later policy can decide mechanically whether that evidence is permitted at inference time. Provenance tells us what role the evidence had; the experiment policy determines whether that role is admissible.


9. What this run establishes, and what it does not

One case, one target, one sequential run per policy. This is a mechanism probe, not the paired repeated design the book uses when it wants an effect estimate. Chapter 6 also explicitly avoided turning observed execution variation into one universal scalar “noise floor.”

The metric defect is the strongest result. The changed calculation is deterministic: punctuation-only changes disappear under token normalization. In this measured ed-003 call, that rule assigned the agent’s punctuation-focused candidate 0.6500. The existence of the scoring failure is therefore established by the metric definition and score decomposition; its frequency in realistic traffic is not.

The four-way tie is weak evidence about retrieval. It says four policy rows received the same scalar score. It does not say the policies supplied equivalent evidence or caused equivalent behavior. Direct context and decision-time case memory actually reached the final rewrite with the same empty evidence string and still produced different text, so at least one visible difference is runtime/order variation rather than retrieval causality.

The guideline corpus is synthetic and tiny. Five hand-written editorial guidelines are not a repository index. Nothing here establishes how these policies scale to thousands of files.

RLM was never exercised successfully. Its selection step failed before exploring the corpus, so the downstream score belongs to the no-evidence fallback path.

The agent result is one trajectory. Relative to fixed retrieval in this run, it used about 9× the total tokens and 9.4× the elapsed time, selected one relevant guideline instead of two, and produced one downstream rewrite. That establishes the observed cost and trajectory, not a general rate or tendency.

What survives all of those limits is structural:

Evidence selection is program behavior. It has a cost, provenance, and failure modes of its own. The downstream evaluator is a separate instrument, and a defect in that instrument can dominate the apparent ranking of evidence policies.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
Different evidence packets receive the same scoreThe metric maps distinct outputs to the same scored featuresCompare evidence packets, output strings, and component breakdownsReport the tie as an evaluator result; do not infer that retrieval had no effect
Identical final inputs produce different rewritesRuntime or execution-order variationCompare the final rewrite input fingerprints, not only policy namesRepeat in balanced order or fresh sessions before attributing the difference to a policy
Adaptive retrieval costs more and scores worseSelection overhead is real, while the score may reflect either evidence quality or evaluator defectsCompare selected IDs, trace, tokens, latency, output, and score components against fixed retrievalMeasure selection and evaluation separately
A labeled-memory policy scores wellIt uses a richer evidence role than decision-time retrievalCheck evidence_role and the deployment evidence policyCompare like with like; block the role only when the protocol declares it inadmissible
An experimental mechanism “matched the baseline”It failed and fell back to no evidenceRead selection_status and the recorded exception, not the scoreReport the failed mechanism; never credit the fallback
Context is huge despite retrievalNo evidence budgetCount characters and tokens per evidence packetBound top-k, line spans, and summaries
The same selector returns different evidence run to runExternal state or iteration order is not pinnedDiff evidence packets across runsSort candidates, fingerprint results, pin corpus revision
Agent retrieval can reach forbidden materialTool surface includes evidence outside the declared experiment boundaryAudit tools, revisions, roles, and split permissionsEnforce the experimental firewall in Chapter 18

Conclusion

We compared six evidence-acquisition policy labels under one final rewrite program and one scoring setup. The resulting rewrite calls did not receive six distinct evidence inputs.

Three rows reached the final rewrite with no supplemental evidence: direct context, decision-time case memory — whose selector returned no case in this run — and the failed RLM path. Direct context and decision-time case memory therefore provide a useful negative control: identical final rewrite inputs produced different strings, so that difference cannot be attributed to retrieved evidence.

Fixed retrieval supplied 337 evidence characters, labeled training memory supplied 138 under a different evidence role, and the agent supplied 204 after spending 3,820 selector tokens. Relative to fixed retrieval, the agent used about 9× the total tokens and 9.4× the elapsed time in this run. It selected the relevant dialogue-punctuation guideline; one trajectory does not establish that adaptive selection was better or worse in general.

The clearest result is in the evaluator. The agent’s punctuation-focused rewrite received 0.6500 because v1 normalizes away punctuation before deciding whether the text changed. Its 1.00 overlap is lexical token recall, not exact punctuation agreement with the human reference. That is enough to establish the scoring defect without claiming that the agent produced the uniquely best rewrite.

RLM supplies a different lesson: its selection mechanism failed before exploring the corpus, and the downstream no-evidence rewrite must not be credited to RLM.

We therefore remove two assumptions without over-reading the single case. Long context does not eliminate the need for an evidence-selection policy, and adaptive selection does not become better merely because an LM chose the evidence. Selection has to be measured with its cost and provenance, while the evaluator has to be audited separately.

The program can now be handed evidence or select it through bounded mechanisms. Chapter 18 enforces which evidence roles, tools, revisions, and cases are admissible inside an experiment.

And Chapter 15’s measured failure is still unaddressed. That agent had the relevant evidence in context and drew the wrong conclusion from it. No retriever fixes that. It requires the program to hold more than one reasoning path open before committing, which is Chapter 17.


Further Reading

  • Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020): the RAG framing this chapter’s fixed-retrieval policy is a small instance of. (arXiv:2005.11401)
  • Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023): why “put it all in the prompt” degrades as the context grows, even when the model nominally supports the length. (arXiv:2307.03172)
  • MemGPT: Towards LLMs as Operating Systems (Packer et al., 2023): treats context as a managed, paged resource rather than a single window — the same instinct behind RLM-style external state. (arXiv:2310.08560)
  • Case-Based Reasoning: Foundational Issues (Aamodt and Plaza, 1994): the retrieve–reuse–revise–retain cycle that “decision-time case memory” and “labeled training memory” are two points on. (AI Communications 7(1))