← DSPy From First Principles

Compile the Program

Treat DSPy compilation as candidate generation under a frozen boundary — then look at what the optimizer actually read, and find that expanding the corpus bought this chapter nothing.

Chapter 9 gave the optimizer an explicit objective and showed a concrete limitation: v1 can assign very high scores to candidates that reverse important parts of the source meaning.

We are going to compile against that frozen metric anyway. Not because the limitation is acceptable, but because changing the objective now would change the experiment. v2 remains outside the main search loop as an independent semantic check.

Compilation in DSPy takes a program, training evidence, a metric, and an optimizer, and returns a candidate program state.

    flowchart LR
    PR[program] --> CO[compile]
    TE[training evidence] --> CO
    ME[metric] --> CO
    OP[optimizer] --> CO
    CO --> CA[candidate program state]
  

This is not compilation in the machine-code sense. The Python architecture stays exactly as written; what changes is the LM-facing state — demonstrations, instructions, and other prompt parameters, depending on the optimizer.

This chapter is about the boundary around that operation, and it deliberately makes no claim about candidate quality. That comes in Chapter 11.

Together, Chapters 10 through 13 preserve the mechanism-level record behind Chapter 9’s negative-control summary: what each optimizer could change, what evidence it consumed, what it cost, and why a successful search-objective result was not sufficient evidence of editorial improvement. Those details matter because the same mechanisms can later be tested in a task with externally executable correctness.

What it does produce is a finding we did not go looking for: expanding the corpus from four to forty-four cases changed the amount of evidence available to compilation without changing how much evidence this BootstrapFewShot configuration actually consumed.

That is not “the larger corpus bought nothing.” The larger corpus is what makes the consumption failure visible.


1. Baseline and candidate are different objects

Name the boundary before crossing it:

EXPERIMENT = {
    "experiment_id": "book-edit-exp-001",
    "baseline_program": "editorial_rewrite_program",
    "baseline_version": "0.1",
    "optimizer": "BootstrapFewShot",
    "optimizer_config": {
        "max_bootstrapped_demos": 2,
        "max_labeled_demos": 0,
        "max_rounds": 1,
    },
    "train_case_ids": TRAIN_IDS,      # 26 cases, 7 families
    "dev_case_ids": DEV_IDS,          # 11 cases, 3 families
    "holdout_case_ids": HOLDOUT_IDS,  # 7 cases, 2 families — not passed to compile
    "metric": "editorial_metric_v1",
}

A candidate is a proposal, not a replacement:

    flowchart TD
    B[baseline v0.1] --> C[compile on allowed data]
    C --> V2[candidate v0.2]
    V2 --> EV[evaluate under the frozen protocol]
    EV --> DEC{accept, reject, or keep investigating?}
  

The candidate is a proposal, not a replacement. The distinction has a name that will recur through chapter 19:

The optimizer proposes. It does not promote itself.

That sounds obvious and is routinely violated, usually by convenience rather than intent — a script that compiles and writes the result straight to the active program path, because the alternative required a second step nobody had time to build.


2. The minimal compile shape

import dspy

def compile_candidate(trainset):
    baseline = EditorialRewriteProgram()
    optimizer = dspy.BootstrapFewShot(
        metric=dspy_editorial_metric_v1,
        max_bootstrapped_demos=2,
        max_labeled_demos=0,
    )
    return optimizer.compile(student=baseline, trainset=trainset)

Note the metric name. Chapter 8 renamed the DSPy-facing wrapper to dspy_editorial_metric_v1 precisely so that this line states, without ambiguity, which version is being maximised. The whole argument of chapters 11 and 12 turns on that being v1.

Optimizers differ in what data they accept. BootstrapFewShot.compile takes student, an optional teacher, and trainset — no validation set. MIPROv2 takes a valset and uses it to select among candidates:

optimizer = dspy.MIPROv2(metric=dspy_editorial_metric_v1, seed=13)
candidate = optimizer.compile(student=baseline, trainset=trainset, valset=devset)

The stable rule underneath the API differences is about declared evidence access, not split names:

train evidence    may be used to construct candidate state
dev evidence      may be used for candidate selection when the protocol allows it
holdout evidence  excluded from optimization for the entire experimental cycle

We ran this with DSPy 3.3.1 on the canonical local Qwen3/Ollama dependency. Compilation made six LM-history entries and consumed 2,025 prompt tokens and 441 completion tokens, 2,466 in total. The returned candidate carried six demonstrations, two per predictor across the three-stage program.

The candidate state itself records which demonstrations were installed, so the important result does not rest on token arithmetic alone: all three predictors received bootstrapped demonstrations sourced from the same first two training examples.

That is compile evidence. It establishes a state transition under the declared boundary. It says nothing about whether the candidate is better.


3. Audit before and after

Compilation should leave a trail, and the trail needs three entries, not two.

def program_state_for_record(program) -> dict:
    if hasattr(program, "dump_state"):
        return program.dump_state()
    return {"class": program.__class__.__name__, "repr": repr(program)}

def compile_with_audit(trainset, devset, holdout):
    baseline = EditorialRewriteProgram()
    before = program_state_for_record(baseline)

    optimizer = dspy.BootstrapFewShot(
        metric=dspy_editorial_metric_v1,
        max_bootstrapped_demos=2,
        max_labeled_demos=0,
    )
    candidate = optimizer.compile(student=baseline, trainset=trainset)

    return {
        "baseline_before": before,
        "baseline_after": program_state_for_record(baseline),
        "candidate_state": program_state_for_record(candidate),
        "visible_train_ids": [e.case_id for e in trainset],
        "reserved_dev_ids": [e.case_id for e in devset],
        "excluded_holdout_ids": [e.case_id for e in holdout],
    }

The baseline-after snapshot is the one people omit, and it is the one that matters. Writing candidate = optimizer.compile(student=baseline, ...) looks like a pure function returning a new object. Nothing in that line guarantees the optimizer did not mutate baseline on the way past. Capturing the baseline’s state on both sides turns that assumption into a checked invariant.

Our measured compile produced three state fingerprints. The first two — baseline before and baseline after — are identical. The third, the candidate, differs. The accepted baseline held zero demonstrations; the candidate held six.

The compile-boundary invariants all held:

  • the baseline state was unchanged by compilation;
  • the candidate was a distinct Python object;
  • only training cases were admissible to this optimizer invocation;
  • the development split was reserved, not passed to compile;
  • the holdout was excluded from the compile boundary;
  • metric, provider, and LM-configuration identities stayed frozen;
  • no promotion occurred.

Notice what is missing: none of those invariants says how much of the admissible training set was actually consumed. optimizer_visible_train_ids records the set passed to compile; it is an access boundary, not a coverage statistic.

Section 4 is what happens when those two concepts are separated.

Compiled DSPy state can be awkward to inspect, which is why tools like dspy-inspector exist. Treat that as a signal rather than a dependency: persist your own candidate state, fingerprints, and run records, and do not bind production contracts to any inspector’s internal view of a framework’s private structures.


4. What the optimizer actually read

Now the part we did not expect.

The training split has 26 cases across 7 families. Chapter 7 built that corpus deliberately because the earlier fixture could not expose enough variation: Chapter 5’s three-case ablation supported a conclusion that reversed on the larger corpus, and Chapter 8’s old baseline reduced evaluation to one development sentence reported with absurd numerical precision.

The question here is different. Having more evidence available does not imply that a particular optimizer configuration will consume it.

So how many of those 26 cases did BootstrapFewShot look at?

Two.

In this measured run, BootstrapFewShot bootstrapped qualifying traces from ed-001 and ed-002 and then reached max_bootstrapped_demos=2. No later training case contributed a bootstrapped demonstration.

The behavior is consistent with the optimizer’s greedy traversal of the supplied trainset: qualifying traces are accumulated until the configured demo quota is filled. For this ordered trainset, the quota was satisfied immediately by the first two examples.

Cases ed-005 through ed-028 — 24 of the 26 training rows, spanning the remaining six families — did not supply demonstrations to this candidate and generated no task-model calls in the compile history after the quota was filled.

That includes the academic-abstract, news-report, and legal-plain-language families. The last of those contains ed-025, the deadline-trigger case that Chapter 4 showed plain Predict getting semantically wrong.

Be precise about the evidence: this run establishes that those rows did not participate in the observed bootstrap trace collection. It does not establish a universal claim about every internal operation BootstrapFewShot could perform under every configuration.

Three consequences worth carrying forward.

List order was part of candidate generation. Nothing in this run established that ed-001 and ed-002 were the most representative or instructive demonstrations. They were the first qualifying examples encountered before the quota filled.

That predicts a useful sensitivity test: shuffle or deliberately permute the training order and ask whether the selected demonstrations and candidate fingerprint change. We have not measured that experiment here, so do not state that a shuffle definitely produces a different candidate.

The engineering lesson does not depend on the unrun test: if greedy traversal can make order decision-relevant, training order belongs in provenance or must be neutralised by the protocol.

The ordinary compile result did not report training coverage as an experimental statistic. The compile completed successfully, returned a distinct candidate, reported token usage, and passed every boundary invariant in Section 3.

Our own audit had recorded all 26 rows as optimizer-visible because they were passed in the trainset. That was correct but incomplete. The candidate state showed six demonstrations sourced from two cases, and the LM history contained six compile calls — three predictors across those two examples. Putting those records together exposed the difference between available evidence and consumed evidence.

The lesson is not that DSPy hid an error. Coverage was simply not one of the quantities our protocol had chosen to measure.

More available data is not a universal fix. This is the one that generalises. Corpus expansion changed the conclusions of earlier chapters because those evaluations actually traversed the larger fixture. Here, the resulting candidate still filled its bootstrap quota after two training examples.

So the larger corpus changed almost nothing about the candidate state, but it changed what we could learn about the optimizer. With only two training cases, a two-example quota looks like full coverage. With twenty-six, the same quota reveals itself as the binding constraint.

When a result will not move, ask two separate questions:

  1. Do we have enough evidence available?
  2. How much of that evidence can this algorithm and configuration actually use?

Collecting more data answers only the first.

The pattern persists into the next chapter. Chapter 11 raises the bootstrap budget to four demonstrations per predictor and its selected demos come from four training cases — ed-001, ed-002, ed-006, and ed-007. That is four of twenty-six training rows, about 15%.

Do not treat that percentage as a universal measure of BootstrapFewShot’s “data usage”; it is a description of the evidence that entered the saved candidate in this run.

Chapter 12’s MIPROv2 has a different search procedure: it proposes and evaluates multiple program states using train and development evidence rather than merely filling a bootstrap quota. That is why optimizer algorithm and budget belong beside dataset size whenever we explain what evidence shaped a candidate.

So before running an optimizer, ask what evidence it is allowed to access and what its budget makes it capable of consuming. Afterwards, inspect what evidence actually entered candidate generation.

A serious optimizer audit separates available evidence, protocol-admissible evidence, and evidence actually consumed by this run.


5. Persist the candidate state

DSPy programs save and load:

candidate.save("artifacts/editorial_rewrite_candidate.json")

loaded = EditorialRewriteProgram()
loaded.load(path="artifacts/editorial_rewrite_candidate.json")

The saved state is necessary and not sufficient. A candidate fingerprint tells you the state differs from the baseline. It does not tell you which optimizer produced it, which examples were visible, which metric selected it, which model performed the compile, or whether holdout evidence leaked into the search.

So write a manifest beside it:

import hashlib
import json
from pathlib import Path

def write_manifest(path: str, payload: dict) -> str:
    blob = json.dumps(payload, indent=2, sort_keys=True)
    fingerprint = hashlib.sha256(blob.encode("utf-8")).hexdigest()
    Path(path).write_text(json.dumps({**payload, "manifest_fingerprint": fingerprint},
                                     indent=2, sort_keys=True))
    return fingerprint

The manifest answers:

Which baseline was compiled?
Which optimizer produced this candidate, with which configuration?
Which case IDs were admissible to it — and which actually contributed to candidate generation?
Which metric, at which version, selected it?
Which model and configuration ran the compile?
Which DSPy version?
Which cases remained untouched?

Section 4 is the reason that line has two halves. Admissibility and consumption are different facts. Recording only the trainset passed to compile can make a 2-of-26 run look like a 26-case optimization.


6. The protocol

Compilation is one step in a sequence that runs from chapter 8 to chapter 19, and it is worth seeing whole:

freeze the manifest              (chapter 8)
audit the corpus and splits      (chapter 7)
evaluate the baseline            (chapter 8)
attack the metric                (chapter 9)
run one declared optimization procedure on its allowed evidence    ← this chapter
persist the candidate
audit candidate integrity
evaluate candidate under the frozen protocol   (chapters 11–13)
apply a conservative promotion policy          (chapter 19)
evaluate the holdout, once                     (chapter 19)
persist the result

Two properties of that list matter more than its contents.

The outer experimental procedure is singular and bounded: one declared optimizer invocation, with a declared internal budget, on declared evidence under a frozen objective.

That does not mean the optimizer itself performs only one trial. MIPROv2 and GEPA later conduct internal search by design; their trial counts are part of the declared optimizer configuration.

The extra search layer appears when we rerun the entire optimization procedure repeatedly and keep the best candidate without recording that best-of-N selection. At that point the outer experiment has become a search over optimizer runs, and any development evidence used to choose the winner has been spent more times than the protocol admits.

And promotion is at the far end, separated from compilation by four chapters of evaluation. That distance is deliberate. DSPy can generate a candidate and report a score for it. It cannot decide that its own candidate is safe to deploy, and any system that lets it has fused proposal with promotion.

In this chapter we do not use the reserved development family to establish candidate quality. Compilation and evaluation are separate experimental events.

That separation is specific to BootstrapFewShot here. Later optimizers such as MIPROv2 explicitly use development evidence inside their search procedure, so their post-compile evaluation has a different interpretation and must be recorded accordingly.


7. What compilation is allowed to change

Decide this before running anything:

ComponentMay change during compile?Reason
Program Python codeNoThis experiment optimizes LM-facing state only
Signature field namesNo in this experimentRenaming a field changes the declared contract
Signature field descriptionsHeld fixed by these optimizer configurationsThey are outside the search surface used in Chapters 10–13
Signature docstrings / instructionsDepends on optimizerMIPROv2 and GEPA mutate instructions; BootstrapFewShot does not
DemonstrationsYesThis is what bootstrap-style optimizers select
MetricNoThe optimizer is being judged by it
Model configurationNoChanging model and program together confounds everything
Holdout casesNoThey are the only untouched evidence you have

The instruction row is the one Chapter 3 warned about. In the optimizer configurations used in this book, field names and descriptions are held outside the search surface while MIPROv2 and GEPA are allowed to mutate signature instructions.

That is a property of the declared optimization surface, not a universal guarantee that descriptions can never change. Another optimizer can choose a wider surface.

So the strongest rule remains:

Anything that must survive regardless of optimization belongs in independently enforced validation or evaluation, not merely in natural-language state that happens to be frozen by today’s optimizer.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
The accepted program changes after compileThe baseline was mutated in placeCompare baseline-before and baseline-after fingerprintsSnapshot both sides; treat compiled output as an artifact
A candidate is returned and improvement is assumedThe compile result was never inspectedCompare candidate and baseline state fingerprintsReport whether state actually changed, and how
Adding training data changes nothingAvailable evidence may exceed what the optimizer budget can consumeCompare admissible case IDs, candidate demo provenance, and compile historyMeasure coverage first; then change budget, ordering policy, or optimizer if justified
Holdout cases influence optimizationThe split boundary was not enforced at the call siteLog the exact case IDs passed to compile and any optimizer-side evaluatorExclude holdout from every optimization/search call
The candidate cannot be reproducedNo manifest beside the saved stateInspect the artifact directoryPersist optimizer, data, model, metric, and protocol fingerprints
Compiling several times and keeping the bestSelection pressure that the protocol does not admitCount compile runs against reported dev evaluationsDeclare the number of runs, or treat it as a search
An optimizer score is treated as deployment approvalProposal and promotion are fusedLook for automatic writes to the active program pathRequire an independent evaluation and policy step

Conclusion

Compilation now happens inside a boundary we can inspect. BootstrapFewShot produced a candidate with six demonstrations and a distinct state fingerprint. The accepted baseline was unchanged. The development split stayed reserved, the holdout stayed excluded and uninspected, and no promotion occurred.

The finding that will outlast this chapter is Section 4. We passed 26 training cases across seven families into the compile boundary. The saved candidate drew its bootstrap demonstrations from two of them, because the configured quota was satisfied immediately.

The larger corpus therefore did two different things. It did not force this candidate to use more training evidence. But it made the mismatch between available evidence and consumed evidence measurable.

That is the general lesson. When a result refuses to move, “collect more data” is only one hypothesis. Before expanding the dataset again, inspect the algorithm, its budget, its traversal policy, and the provenance of the evidence that actually shaped the candidate.

We removed the assumption that an optimizer’s output should replace the accepted program, and the assumption that calling compile is itself evidence that something improved. This chapter establishes a state transition under a controlled boundary. It establishes nothing about quality.

Which brings us to the question we have been deferring since Chapter 8. We have a baseline whose reported development runs sit around 0.79, a candidate carrying six demonstrations bootstrapped from two training cases, and a v1 objective that Chapter 9 showed can score specific semantic reversals highly.

Does the candidate actually help?