← DSPy From First Principles

Few-Shot Optimization

Compile a few-shot candidate, watch it improve one metric and damage the other, and find out why compiling against the better metric changes nothing at all.

Chapter 10 produced a candidate and refused to say whether it was any good. This chapter says.

The mechanism is few-shot optimization, and the question is practical:

Can we improve the program by selecting better demonstrations, rather than rewriting the program ourselves?

The answer, on our fixture, is the most interesting result in this book. In one run, the candidate appears to improve the metric we optimized against while damaging the metric we did not. Under seven paired fresh-process sessions, the apparent v1 gain collapses to essentially zero while the v2 loss remains large and systematic.

Then we try the obvious repair: compile against v2 instead. The resulting candidate is byte-for-byte identical to the v1-compiled candidate.

So this chapter gives us three different facts that must not be collapsed into one score: a noisy aggregate gain, a reproducible semantic regression, and an objective change that has no effect on the candidate because the relevant failure never enters this optimizer’s compile loop.


1. What BootstrapFewShot does, and what its metric sees

BootstrapFewShot builds demonstrations for each predictor in a program. Some come from labeled examples; the interesting ones are bootstrapped — produced by running the program on a training case and keeping the trace only if the metric approves.

    flowchart TD
    TE[training examples] --> AT[student attempts each case]
    AT --> CT[candidate traces]
    CT --> MF{metric approves?}
    MF -->|yes| DEM[trace becomes a demonstration]
    MF -->|no| DROP[discarded]
    DEM --> CS[compiled student]
  

The filter is the whole mechanism:

Only traces that satisfy the metric become behavior guidance.

Chapter 9 already told us what that implies here. v1 can assign very high scores — about 0.85 to 0.96 on the three canonical attacks — to candidates that reverse important parts of the source meaning.

A bootstrapped demonstration is therefore evidence that a trace passed the compile metric. It is not independent evidence that the trace is semantically safe, representative of the target domain, or acceptable to a human reviewer.

But notice something more specific about where that metric is applied, because it becomes the point of section 7. The filter runs over training traces only. The metric never sees a development case during compilation. BootstrapFewShot has no validation set — chapter 10 noted that compile takes student, teacher, and trainset, and nothing else.

So the compile metric’s admissible field of view is the 26 training cases. The actual field of view of this run is narrower still: five training traces are scored, four clear the threshold, and those four become the source cases for the saved demonstrations.


2. Compiling the candidate

import dspy

def compile_few_shot_candidate(trainset):
    student = EditorialRewriteProgram()
    optimizer = dspy.BootstrapFewShot(
        metric=dspy_editorial_metric_v1,
        metric_threshold=0.7,
        max_bootstrapped_demos=4,
        max_labeled_demos=0,
        max_rounds=1,
    )
    return optimizer.compile(student=student, trainset=trainset)

Then evaluate both programs under the frozen protocol from chapter 8 — same model, same cases, same metric, holdout untouched:

baseline = EditorialRewriteProgram()
candidate = compile_few_shot_candidate(trainset)

evaluate_dev = dspy.Evaluate(devset=devset, metric=dspy_editorial_metric_v1)
baseline_dev = evaluate_dev(baseline)
candidate_dev = evaluate_dev(candidate)

The original measured v1 compile took 30.9 seconds and 6,195 task-model tokens. The metric was called five times during bootstrapping; four traces cleared the 0.7 threshold.

Keep both counts. Five traces were inspected; four were admitted. Candidate provenance and optimizer-consumption provenance are not the same statistic.


3. What it selected

Four accepted traces, distributed across three predictors:

PredictorDemonstrations
analyze4
rewrite4
assess4

Twelve demonstrations in total, all bootstrapped, all traceable to one of four training cases: ed-001, ed-002, ed-006, ed-007.

Four cases out of twenty-six supplied the saved demonstrations, but the optimizer inspected five training traces to get them.

In canonical train order those are ed-001, ed-002, ed-005, ed-006, and ed-007. Four cleared the threshold: ed-001, ed-002, ed-006, and ed-007. ed-005 was evaluated and rejected.

That gives the consumption funnel Chapter 10 asked us to record:

  • 26 training cases were admissible;
  • 5 were inspected by the compile metric;
  • 4 supplied demonstrations;
  • 21 were never reached in this compile;
  • 22 supplied no demonstration to the candidate.

Raising max_bootstrapped_demos from 2 to 4 therefore expanded the inspected prefix, but it still left most of the available training evidence outside candidate construction.

The provenance checks passed cleanly. No development case became a demonstration. No holdout case appeared in the compile evidence. Every saved demonstration resolved to a known training case with a complete trace.

We can therefore distinguish three questions precisely:

  1. Was the evidence admissible? Yes.
  2. Which evidence was inspected and selected? Five inspected, four selected.
  3. Were the selected demonstrations good guidance for every development family? That remains an empirical question.

Split correctness is necessary. It is not demonstration quality.


4. The result

First, preserve the original single-run result because it exposes the concrete mechanism. On that one 11-case development evaluation, baseline and candidate scored:

ProgramMetric v1Metric v2
Frozen baseline0.78930.7893
Few-shot candidate0.80050.7399
Delta+0.0112−0.0494

On that measurement occasion, the candidate is higher under the metric used for compilation and lower under the semantic guardrail.

Do not promote those deltas into final effect estimates yet. Section 6 repeats the comparison across seven fresh-process paired sessions and changes the interpretation of the positive number completely.

What this single run is excellent for is diagnosis: it gives us exact cases and exact sentences to inspect.

The per-case breakdown shows where:

Casev1v2Note
ed-0320.782 → 0.8180.782 → 0.818higher in this run
ed-0330.818 → 0.9270.818 → 0.927higher in this run; known session-sensitive case
ed-0360.956 → 1.0000.956 → 1.000higher in this run
ed-0030.877 → 0.8770.877 → 0.877canonical case, unchanged
ed-0351.000 → 0.9671.000 → 0.300qualifier dropped; semantic violation
ed-0381.000 → 0.9671.000 → 0.967small structural loss
ed-029, ed-030, ed-031, ed-034, ed-0370.650 → 0.6500.650 → 0.650unchanged at the structural floor

Three cases move upward in this run. Five stay pinned at 0.650. ed-038 loses 0.033 under both metrics. And ed-035 loses only 0.033 under v1 but 0.700 under v2.

That is why the two metrics disagree so sharply. v1 treats the ed-035 edit as one small lexical loss among several ordinary case movements. v2 treats it as a disqualifying semantic event.


5. ed-035

ed-035 belongs to the marketing-copy family. The canonical source sentence is:

Using the filter regularly can help reduce limescale build-up in your kettle.

The editorial goal is Make the benefit statement more direct. The context says this is a regulated category and the claim must remain at helps reduce, not prevents.

The reference rewrite is:

Using the filter regularly helps reduce limescale build-up in your kettle.

Two semantic constraints are recorded with the case:

  • keep the claim at helps reduce, not prevents or eliminates;
  • the benefit depends on regular use.

The frozen baseline produced the reference rewrite exactly and scored 1.000. The few-shot candidate returned:

Using the filter helps reduce limescale build-up in your kettle.

One word removed. regularly.

As line editing, the cut is plausible. regularly can look like expendable modifier language, and the remaining sentence is cleaner and fully grammatical. This is not a model producing nonsense.

It is also worth using the case’s actual goal. The model was asked to make the benefit statement more direct, not generically to maximize concision. Removing can from the source into the reference satisfies that goal while preserving the condition. Removing regularly goes one step further and changes what the claim depends on.

But regularly is not decoration in this fixture. It carries the condition named explicitly by the semantic constraint.

With it, the claim is about the benefit of regular use. Without it, the sentence attributes the benefit to use of the filter without preserving that condition.

We do not need a jurisdiction-level legal claim to establish the regression. The candidate violates the frozen task requirement on its own terms.

The metrics disagree about how much that matters.

v1 charged 0.033. The candidate still shares almost all of its vocabulary with the reference, still changed something, still stayed in scope. Deleting one adverb barely registers as a lexical event, so the score falls from 1.000 to 0.967 — which reads, in a results table, as noise.

v2 charged 0.700. The judge evaluated the candidate against the recorded regular-use constraint, returned violated with reason code meaning_weakened, and the violation cap took the score from a structural 0.967 to 0.300.

As Chapter 9 established, the verdict is evidence from a validated but same-model semantic guardrail rather than an independent authority. Here the disputed text and the constraint are simple enough to inspect directly.

That is closely related to ed-025 in Chapter 4, where Predict changed a deadline trigger from delivery to receiving the item.

The surface operations differ — ed-035 drops a scope-bearing word, while ed-025 substitutes one event description for another — but the metric failure is the same: a lexically small edit changes a condition or scope that the structural score cannot represent.

Both examples are dangerous precisely because the resulting prose remains fluent and plausible.

Now recall where the saved demonstrations came from. The four source cases — ed-001, ed-002, ed-006, and ed-007 — belong only to chapter-opening and close-third-narration.

No marketing, legal, technical, news, or instructional family contributes a demonstration to this candidate. That concentration gives us a plausible mechanism for cross-domain style transfer: examples that reward tightening in literary prose may bias later rewrites toward similar compression.

But we did not run a demonstration-family ablation, so do not claim that those four demos caused the deletion of regularly. The measured claim is narrower: a candidate built from demonstrations concentrated in two literary families produces the ed-035 regression on an unseen marketing family.

Chapter 10’s coverage finding and this chapter’s failure therefore connect without requiring a causal story we have not tested.


6. Is the gain even real?

The single-run +0.0112 result is no longer the right evidence for the aggregate effect. We repeated the comparison in seven fresh-process sessions, pairing candidate and baseline case by case and rotating execution order.

The result:

Paired candidate − baseline deltaMeanSDRangeInterpretation
v1+0.00030.0016−0.0020 to +0.0013no measurable aggregate improvement
v2−0.06030.0016tightly negativelarge systematic regression

The v1 candidate’s aggregate delta is positive in 5 of the 7 sessions and negative in the other two, and the mean sits well inside the paired SD. The original +0.0112 was therefore not a small effect hiding near a heuristic noise threshold; it was a measurement-occasion result that collapses under the stronger paired design.

ed-033 explains why the earlier run was especially vulnerable to this mistake. Across earlier sessions, the unchanged baseline itself can land on alternative rewrites for that case. Within the seven-session paired batch the baseline happens to be stable, but pairing prevents that historical session effect from being misread as a candidate effect.

The semantic result behaves differently. The candidate drives ed-035 to v2 0.300 in every run, and the paired aggregate v2 loss is about thirty-eight times the observed paired SD.

So the defensible statement is:

The few-shot candidate has no measured aggregate advantage under v1 in the paired experiment, while it introduces a reproducible semantic regression that v1 prices as a minor lexical change.

That is the reason this book insists on separating candidate generation from promotion. If we had stopped after the first run, reported +0.0112, and promoted on sign alone, we would have had a numerical justification for accepting a candidate that reproducibly drops a condition from ed-035.

The failure is not that the first number was fabricated. It was a real measurement. The failure would have been treating one measurement occasion as sufficient evidence for replacement.


7. “Then why not compile against v2?”

This is the obvious objection, and we ran the experiment.

The setup: the same trainset, the same optimizer, the same threshold and demo budget, but with the compile metric changed from v1 to v2.

If the regression were caused by v1 admitting a training trace that v2 would reject, this intervention should change the selected demonstrations. That is the mechanism the experiment actually tests.

The result:

The two compile arms produced identical candidate state. Not merely the same demo IDs: the saved candidate files are byte-identical and share the same Git blob.

Both contain twelve demonstrations, sourced from the same four accepted cases and distributed identically across the three predictors.

The objective change did have a cost. In the controlled arm, v1 compilation took about 26.4 seconds, made 5 metric calls and no judge calls; v2 compilation took about 41.5 seconds, made the same 5 metric calls plus 12 semantic-judge calls, while consuming the same 6,195 task-model tokens.

The extra semantic evaluation changed the bill but not the candidate.

In the original one-shot dev evaluation the two byte-identical candidates naturally produced the same scores. In the later interleaved seven-session experiment they show tiny score differences despite identical saved state. That is a useful negative control: those differences are measurement-order/runtime variation, not program differences.

ed-035 still breaks under the candidate state.

The explanation is Section 1. BootstrapFewShot applies its compile metric to training traces. Five traces were scored in each arm; the same four cleared the 0.70 threshold under v1 and v2.

The semantic judge therefore had twelve per-constraint opportunities to alter the v2 arm’s accept/reject decisions and changed none of them. The accepted demonstration set remained identical.

And ed-035 is a development case. It never enters the compile loop at all. The compile metric was never asked about it, under either version, because BootstrapFewShot has no mechanism for asking.

compile may access:       26 training cases
traces actually scored:     5
traces accepted as demos:   4
regression observed on:      development case ed-035
ed-035 compile exposure:     none

The lesson is narrower than a universal claim about compile metrics:

Changing only the compile metric from v1 to v2 did not close this regression under this corpus and BootstrapFewShot configuration.

Why? ed-035 is development-side evidence and never enters this compile loop, while the training-trace decisions that do enter the loop are unchanged under v2.

A different training corpus, ordering, threshold, demo budget, optimizer, or metric could change candidate generation. Chapter 12 matters precisely because MIPROv2 uses development evidence during search, so v2 would have a different opportunity to act there.

The practical consequence is that a better metric is not a substitute for evaluating outside the loop. Whatever you compile against, something independent has to look at what came out.


8. Labeled demonstrations and bootstrapped demonstrations

Worth being precise about the two kinds, because they carry different risks.

A labeled demonstration comes from the dataset. It shows the reference output:

Here is an accepted rewrite.

A bootstrapped demonstration is generated by running the program and keeping the trace if the metric approves:

Here is a rewrite this program produced that passed our metric.

The second statement inherits every weakness of the metric used to admit the trace. Four traces scored above 0.7 and became behavior guidance for all three predictors. The resulting candidate later drops a scope-bearing condition on ed-035.

We have not isolated which demonstration, if any, caused that behavior, so “the demos taught the program to delete qualifiers” would overstate the evidence.

The defensible principle is still important: metric design governs both which candidate wins and, for bootstrap-style optimization, which generated traces are allowed to become future context. Admission therefore deserves its own audit, because a demonstration does not announce which behaviors a downstream case will imitate.

One more asymmetry. In the measured run, ed-001 scored 0.92 and ed-002 about 0.964 during bootstrapping — both comfortably above threshold. A threshold tells you a trace was good enough. It does not tell you the trace was representative, and with demonstrations selected by list order, nothing in this pipeline is asking that question.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
The compile metric improves and users complainOptimized against a proxy the users do not shareScore the candidate under an independent metricEvaluate outside the compile loop, always
A small gain is reported as an improvementOne execution schedule is being treated as an effect estimateRun paired comparisons across fresh processes with rotated orderReport the paired delta distribution and sign counts
The gain comes from one unstable caseAggregate movement was never localizedBreak the result down case by case and compare repeated outputsSeparate candidate effects from session/order effects
Upgrading the compile metric changes nothingThe regression is outside the metric’s field of viewCheck which split the compile metric evaluatesAdd independent evaluation; do not expect the filter to catch it
Demonstrations come from one domainSelection is by list order, not diversityGroup demo sources by familyShuffle, stratify, or select deliberately
The candidate learns a habit nobody asked forDemonstrations encode an accidental styleRead the selected demonstrations end to endDiversify sources or reduce the demo count
Prompt cost grows with no quality gainDemonstrations consume context without earning itCompare baseline and candidate token usage per caseUse fewer demos, or reject the candidate
A holdout case appears in the demosSplit boundary violated at the call siteCompare demo source IDs against holdout IDsRebuild the split and rerun from scratch

Conclusion

BootstrapFewShot inspected five training traces, admitted four above the 0.7 threshold, and turned those four traces into twelve demonstrations across three predictors.

The original one-shot evaluation reported +0.0112 under v1 and −0.0494 under v2. The seven-session paired experiment gives the result we should actually carry forward: +0.0003 mean paired delta under v1 and −0.0603 under v2, with paired SD about 0.0016.

So the apparent v1 improvement disappears under repetition. The semantic regression does not. ed-035 reaches v2 0.300 in every run after the candidate deletes regularly. On that case v1 prices the deletion at only 0.033; v2 prices the violated constraint at 0.700.

Then we tried changing the compile metric. The v2 arm made 12 additional judge calls and returned a byte-identical candidate, because the same four training traces cleared the threshold and ed-035 never enters BootstrapFewShot’s compile loop.

The important lesson is not that a better compile metric is powerless. It is that an objective can only change search through evidence the search procedure actually consults.

We removed the assumption that adding accepted demonstrations is automatically an improvement. And we removed something larger: the assumption that improving the thing you optimize against is sufficient protection against optimizing the wrong thing. It is not, because the optimizer and your evaluation do not necessarily look at the same evidence.

What we have not yet tested is whether this is a property of BootstrapFewShot or of optimization generally. This optimizer only selects demonstrations — it never touches an instruction, and it never uses the development set for anything.

The next one does both.

What happens when the optimizer can rewrite the program’s instructions and select against the development set directly?