Optimize the Instructions
Let MIPROv2 search over instructions and demonstrations, decompose the improvement it finds down to the single case that produced it, and watch a second optimizer break the same sentence as the first.
Chapter 11’s optimizer could only select demonstrations. It never touched an instruction, and it never looked at the development set — which is why upgrading its metric changed nothing.
MIPROv2 does both.
instruction candidates
+
demonstration candidates
+
metric
+
development data
↓
candidate program
This is a strictly larger search over a strictly larger space, evaluated against evidence the previous optimizer could not see. If chapter 11’s result was an artifact of a weak mechanism, this is where it gets corrected.
It does get corrected, partly. MIPROv2 finds a real improvement that chapter 11 could not. It also breaks the same sentence, in the same way, for the same reason.
1. What MIPROv2 searches
MIPROv2 generates few-shot examples and new instructions for each predictor, then searches combinations of them under a budget:
flowchart TD
BE[bootstrap few-shot examples] --> PI[propose instruction candidates]
PI --> EC[evaluate combinations against the validation set]
EC --> B{budget left?}
B -->|yes| PI
B -->|no| BEST[return the best state found]
Instruction candidates are proposed by a language model, informed by summaries of the training data, the program’s structure, the bootstrapped examples, and sampled generation tips. This means a second model role appears in the pipeline — the prompt model, which writes instructions, distinct from the task model, which executes them. Chapter 6’s provenance table warned that these should be recorded separately, and this is where that matters: if the prompt model changes, the search path changed even though nothing else did.
The consequential difference from chapter 11 is the validation set. MIPROv2.compile accepts a valset and evaluates every candidate combination against it. So unlike BootstrapFewShot, this optimizer’s metric does see development cases — repeatedly, by design.
Hold on to that. It changes what the development score means in section 6, and it explains why the v2 result here differs in character from chapter 11’s.
2. The compile shape and the budget
import dspy
def compile_mipro_candidate(trainset, devset):
student = EditorialRewriteProgram()
optimizer = dspy.MIPROv2(
metric=dspy_editorial_metric_v1,
auto=None,
num_candidates=2,
seed=13,
metric_threshold=0.7,
max_bootstrapped_demos=2,
max_labeled_demos=0,
)
return optimizer.compile(
student,
trainset=trainset,
valset=devset,
num_trials=4,
minibatch=False,
)
Note auto=None. DSPy’s default is auto="light", which chooses a budget for you; setting it to None and specifying num_candidates and num_trials explicitly means the search size is a recorded experimental parameter rather than a framework default that may change between versions.
That is not a stylistic preference. A candidate found by a heavy search and a candidate found by a small one are different artifacts, and a manifest that says auto="light" records the name of a policy rather than the size of a search.
The data boundary is unchanged: 26 training cases, 11 development cases as the validation set, and the holdout excluded and uninspected.
MIPRO_CONFIG = {
"optimizer": "MIPROv2",
"auto": None,
"num_candidates": 2,
"num_trials": 4,
"seed": 13,
"minibatch": False,
"metric": "editorial_metric_v1",
"train_case_ids": TRAIN_IDS, # ed-001 … ed-028, 7 families
"dev_case_ids": DEV_IDS, # ed-003, ed-029 … ed-038, 3 families
"holdout_case_ids": HOLDOUT_IDS, # excluded from compile
}
3. What the search returned
MIPROv2 selected a changed candidate. It rewrote the instructions on the analyze and rewrite predictors and attached three bootstrapped demonstrations sourced from ed-020.
Start with the single 11-case development evaluation, the same way chapter 11 did:
| Program | Metric v1 | Metric v2 |
|---|---|---|
| Frozen baseline | 0.7893 | 0.7893 |
| MIPROv2 candidate | 0.8192 | 0.7586 |
| Single-run delta | +0.0298 | −0.0308 |
The same shape as chapter 11, with larger numbers in both directions. Up on the metric we searched against, down on the metric that can see meaning. And, exactly as in chapter 11, a single run is not yet an effect estimate — section 4 repeats it across seven paired sessions.
The search cost, for the record:
| Search resource | Measured |
|---|---|
| Elapsed time | 404.8 s |
| Prompt-model history entries | 4 |
| Prompt-model tokens | 6,523 |
| Task-model history entries | 168 |
| Task-model tokens | 79,084 |
| Optimizer metric calls | 56 |
Roughly 85,600 tokens and seven minutes to move a development mean by 0.03 on eleven cases. Whether that is expensive depends entirely on what the 0.03 turns out to be — which is section 4.
4. Decomposing the gain
The +0.0298 single-run delta clears chapter 6’s ±0.015 noise floor, so unlike chapter 11’s gain it is not obviously noise. But an aggregate is a summary, and chapter 8’s lesson was that the summary is the least informative thing you compute. So break it apart.
Five cases moved. Six did not.
| Case | Baseline | Candidate | Change | Contribution to the mean |
|---|---|---|---|---|
ed-029 | 0.650 | 1.000 | +0.350 | +0.0318 |
ed-033 | 0.818 | 0.927 | +0.109 | +0.0099 |
ed-035 | 1.000 | 0.967 | −0.033 | −0.0030 |
ed-003 | 0.877 | 0.846 | −0.031 | −0.0028 |
ed-038 | 1.000 | 0.933 | −0.067 | −0.0061 |
| Total | +0.0298 |
The arithmetic closes exactly, and it says something the aggregate hid completely.
ed-029 alone contributes +0.0318 — more than the entire reported gain. Everything else the optimizer did nets out to −0.0020. Four of the five cases that moved got worse.
So what happened on ed-029? It is a dialogue case with a target_failure of invalid_output, and the baseline has never once fixed it — it sits at exactly 0.650, the do-nothing floor, in every run we have made across every chapter. MIPRO’s rewritten instruction told the program to preserve all key information, and that was enough. The case went to a perfect 1.000.
That is a genuine, reproducible, mechanistic improvement. An optimizer proposed an instruction a human had not written, and it fixed a case a human-written instruction never fixed. This is what instruction optimization is supposed to do, and it worked.
The other apparent gain is less trustworthy. ed-033 is the coffee sentence from chapter 6 with two attractor rewrites — the baseline produces both across sessions, worth 0.818 or 0.927 depending on the day. Its +0.0099 contribution is largely the same coin flip that inflated chapter 11’s result.
Strip that out and the honest summary is: MIPROv2 fixed one case, made three slightly worse, and got lucky on a fourth.
That is a real result. It is also a much smaller result than “+0.03 improvement” suggests, and you cannot see it without the per-case table.
Across seven paired fresh-process sessions, the aggregate v1 delta is +0.0199, identical in every session (SD 0.0000), candidate ahead in 7 of 7. Unlike chapter 11’s bootstrap gain, this one does not collapse under pairing — it is real and reproducible. It is simply smaller than the single run, which caught ed-033 on its higher rewrite. The paired v2 delta is −0.0407, also identical across sessions: the candidate that raised the objective lowered the independent check every time. +0.0199 is the number to carry forward.
5. ed-035, again
Now the row in that table we have not discussed.
ed-035 went from 1.000 to 0.967 under v1 — a −0.0030 contribution, invisible in the aggregate, indistinguishable from rounding.
Under v2 it went from 1.000 to 0.300.
It is the same case chapter 11 broke, and it broke the same way. The candidate deleted regularly from the water-filter sentence, turning a claim conditional on regular use into an unconditional product claim. The judge returned violated with reason scope_changed. The cap applied.
Run the same arithmetic as chapter 4:
ed-035 penalty under v1: 0.033
ed-035 penalty under v2: 0.700
additional v2 penalty: 0.667
spread across 11 cases: 0.667 / 11 = 0.0606
v1 delta − v2 delta: 0.0298 − (−0.0308) = 0.0606
(paired: 0.0199 − (−0.0407) = 0.0606)
The entire divergence between the two metrics is one sentence. Exactly as in chapter 4, where ed-025’s delivery-versus-receipt substitution accounted for the whole gap, to four decimal places.
And now consider what makes this the most important result in the book.
Chapter 11’s optimizer selected demonstrations from four literary training cases. Chapter 12’s optimizer rewrote instructions, proposed by a language model, informed by a summary of all 26 training cases, and selected against 11 development cases through 56 metric evaluations.
Different mechanisms. Different evidence. Different search spaces. Different amounts of compute, by an order of magnitude.
Both produced a candidate that deletes regularly from ed-035.
That is no longer a quirk of one optimizer. It is what maximising this metric does. Two independent search processes, given the same objective, found the same exploit — because the exploit is not in the optimizer, it is in the objective. v1 pays for concision and cannot see conditionality, so both searches learned to remove a qualifier, and both were rewarded for it.
There is an irony worth pausing on. The instruction MIPRO chose — the one that fixed ed-029 — told the program to preserve all key information. That same candidate then dropped the key qualifier from a regulated claim. The instruction was not wrong. Nothing in the loop could tell it which information was key, because the only signal it had was a score that did not know either.
6. Development data is not holdout data
MIPROv2 evaluated candidate combinations against the development set 56 times. Those eleven cases are no longer neutral evidence about the candidate; they helped choose it.
That is not misconduct. It is exactly what a development set is for. But it changes the status of the number:
train: build candidate ingredients
dev: select among candidate combinations ← used 56 times here
holdout: test the selected candidate, once, after search stops
The candidate’s 0.8192 is selection evidence. It is the score of the configuration that scored best out of everything tried, on the data used to try them, which is a biased estimate by construction — the more you search, the more of the development set’s noise you fit.
Chapter 11’s candidate is in a different position: BootstrapFewShot never touched the development set, so its 0.8005 is closer to an honest out-of-search estimate. The two numbers are not directly comparable, even though they came from the same protocol.
This is why “we did not train on the test set” is an insufficient defence. If a search process selected a candidate after observing performance on a set, that set has become development data regardless of what the directory is called. The only cases that retain independent status are the ones nothing has looked at, and there are seven of them, still sealed.
7. Three stages, one pattern
| Program | Metric v1 | Metric v2 | Paired v1 Δ | What changed |
|---|---|---|---|---|
| Frozen baseline | 0.7893 | 0.7893 | — | — |
| BootstrapFewShot candidate | 0.8005 | 0.7399 | +0.0003 | 12 demonstrations, 4 training cases |
| MIPROv2 candidate | 0.8192 | 0.7586 | +0.0199 | 2 rewritten instructions, 3 demonstrations |
(The v1 and v2 columns are one representative session; the paired column is the seven-session mean.)
Both optimizers improved the objective. Both degraded the independent check. Both broke ed-035.
Neither clears the 0.05 promotion margin that chapter 19 will apply — bootstrap’s paired gain is inside the noise band and MIPRO’s +0.0199 is well under the margin — so neither is a promotion candidate under this book’s policy, and that is before considering that both are worse under v2 than the program they were meant to improve.
The chemistry TPSA project is a useful contrast here. It runs MIPRO inside a comparable protocol — split data, saved programs, ablations — but its target is a physical quantity with an unambiguous error measure. When the metric is MAE against a known value, none of this chapter’s ambiguity exists: a lower error is a better program, full stop. Editorial quality has no such anchor, which is why the metric needs attacking and the results need decomposing.
The general lesson is not that optimizers are untrustworthy. It is that an optimizer is a faithful amplifier of your metric, and the more effective it is, the more precisely it will express whatever your metric got wrong. Chapter 11 amplified a blind spot with 6,195 tokens. Chapter 12 amplified the same blind spot with 85,600.
8. Persist the artifact
candidate.save("artifacts/editorial_mipro_candidate.json")
write_manifest(
"artifacts/editorial_mipro_manifest.json",
{
"program": "editorial_rewrite_program",
"candidate_version": "0.3",
"optimizer": "MIPROv2",
"optimizer_config": MIPRO_CONFIG,
"metric": "editorial_metric_v1",
"instructions_changed": ["analyze", "rewrite"],
"demo_source_case_ids": ["ed-020"],
"dev_metric_calls": 56,
"search_tokens": 85_607,
"model": "recorded_at_runtime",
"prompt_model": "recorded_at_runtime",
"dspy_version": "recorded_at_runtime",
},
)
Two fields here that chapter 10’s manifest did not need. instructions_changed records which predictors the optimizer rewrote, because a candidate that changed instructions is a different kind of artifact from one that only added demonstrations — chapter 3’s durability warning has now actually happened, and the manifest should say so. And dev_metric_calls records how hard the development set was used, which is the number that determines how much to discount the selection score in section 6.
The question the artifact must answer is unchanged:
Could someone rerun or audit this experiment later?
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| A gain clears the noise floor and still misleads | The aggregate hides offsetting per-case movement | Decompose the delta case by case and check it sums | Report which cases moved, not just the mean |
| The whole improvement is one case | Genuine, but far narrower than the headline suggests | Compute each case’s contribution to the mean | State the finding at the resolution the data supports |
| Dev score rises, holdout falls | The search fit development noise | Compare failure taxonomies across splits | Reduce search budget, add cases, or discount the selection score |
| Two optimizers make the same mistake | The exploit is in the metric, not the optimizer | Look for a case both candidates damage | Fix the objective; a different search will not help |
| An instruction says the right thing and the program does the wrong thing | Nothing in the loop can identify what “key” means | Check what signal the instruction is graded against | Encode the property in the metric, not the prose |
| Search consumes far more compute than expected | Budget was left to the framework default | Inspect auto, num_candidates, num_trials in the manifest | Set the budget explicitly and record it |
| The candidate cannot be reproduced | Instructions and demos saved without a manifest | Inspect the artifact directory | Save state plus optimizer, data, model, and metric identity |
| A constraint written in a docstring disappears | The optimizer rewrote the instruction containing it | Diff candidate instructions against the originals | Move durable constraints into field descriptions or the metric |
Conclusion
MIPROv2 searched instruction and demonstration combinations for seven minutes and 85,600 tokens, and returned a changed candidate: two rewritten instructions, three demonstrations, and a development score of 0.8192 against the baseline’s 0.7893 — a paired v1 gain of +0.0199, identical across all seven sessions.
Decomposed, that gain — a reproducible +0.0199 across seven paired sessions — is one case. ed-029 — an invalid_output case the baseline has never fixed in any run of any chapter — went from the 0.650 floor to a perfect 1.000, contributing +0.0318 on its own. Everything else the optimizer did netted −0.0020, and one of the apparent gains was the unstable coffee sentence flipping the way it sometimes does anyway.
So the honest report is: a language model wrote an instruction that fixed a case humans had not fixed, and made three other cases slightly worse doing it. That is a real result, and it is a much narrower one than the headline number.
Under v2 the candidate scores 0.7586, below the baseline. The entire 0.0606 gap between the two metrics’ verdicts is ed-035 — the same sentence chapter 11 broke, damaged the same way, by an optimizer with a completely different mechanism.
That recurrence is the finding. One optimizer selecting demonstrations from four literary cases, and one optimizer writing instructions with a language model after 56 evaluations against the development set, independently arrived at the same edit. The exploit is not in either search. It is in the objective they share.
We removed the assumption that human-authored instructions are fixed, and the assumption that a larger search produces a more trustworthy result. A larger search produces a more thorough exploration of whatever you told it to want.
What remains is that both optimizers received a single number per candidate and no explanation of it. ed-035 scored 0.967 under v1; neither optimizer was told which word was the problem, because a scalar cannot say.
What changes when the optimizer receives feedback in language rather than a score?