Separate What From How
Hold the task contract completely fixed, change only the execution strategy, and measure what the extra reasoning actually bought across 222 local calls.
Chapters 2 and 3 built a contract and said nothing about how a model should satisfy it. That was deliberate, and this chapter collects the payment.
Three objects must remain distinct throughout the comparison:
task contract what valid behavior must preserve
execution strategy how the model attempts the task
evaluation how observed behavior is judged
The signature holds the first fixed. Predict and ChainOfThought vary the second. The frozen fixture and metrics supply the third. If two of those change together, the result no longer tells us what the reasoning policy caused.
SAME CONTRACT
input ────────> Signature ────────> output
│
execution policy
/ \
Predict ChainOfThought
Because the task-facing contract is fixed, swapping the execution strategy gives us one declared experimental factor: Predict versus ChainOfThought. The strategy change will have consequences — a different rendered prompt, a reasoning field, different token use, different latency — but those are effects of the intervention rather than extra knobs changed by hand.
That is rare enough in language-model work to be worth stopping on. Many comparisons between prompting strategies change the instruction, the examples, the output format, and the strategy simultaneously, and then report which bundle felt better.
This chapter runs the comparison properly — 222 measured calls on a local model — and gets a result that surprised us, for a reason that turns out to be the subject of the rest of the book.
1. Two policies, one contract
Direct prediction:
import dspy
class RewriteSentence(dspy.Signature):
"""Rewrite one sentence to satisfy an editorial goal
while preserving meaning, entities, and voice."""
sentence: str = dspy.InputField(desc="The exact sentence to rewrite")
goal: str = dspy.InputField(desc="The local reason this sentence is being edited")
context: str = dspy.InputField(desc="Nearby prose needed to preserve continuity and voice")
rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
rationale: str = dspy.OutputField(desc="Short explanation of the edit")
direct = dspy.Predict(RewriteSentence)
Explicit reasoning, same signature:
reasoned = dspy.ChainOfThought(RewriteSentence)
ChainOfThought extends the signature with an additional reasoning output that the model fills before producing the declared fields. The contract does not change:
sentence, goal, context → rewritten_text, rationale
The policy does:
Predict: answer directly
ChainOfThought: produce intermediate reasoning, then answer
So the calling code is identical for both, which is the property that makes the experiment possible:
def run_case(program, case):
return program(
sentence=case["sentence"],
goal=case["goal"],
context=case.get("context", ""),
)
One implementation note. The shared repository RewriteSentence used by the measured runner already carries two additional inputs, issue_summary and preservation_notes, which Chapter 5 uses when the program is decomposed. The listing above shows the task-facing core for clarity. In this experiment both additional fields are passed as empty strings to both strategies.
So Predict and ChainOfThought receive the exact same expanded signature and the exact same values. The only deliberate intervention is the execution policy.
2. Designing a comparison that means something
The experiment is worth describing before the results, because the design is most of the value.
| Held fixed | Varied |
|---|---|
| Signature and field descriptions | The module: Predict or ChainOfThought |
| Model, quantization, and generation settings | |
| Cases and repetition count | |
| Balanced execution-order policy | |
| v1 and v2 metric definitions | |
| Semantic-judge configuration for v2 |
The model was Qwen3 8.2B, Q4_K_M, served through Ollama and called via DSPy 3.3.1 at temperature 0. Qwen3’s native thinking mode was disabled, so that model-level reasoning would not contaminate the distinction between DSPy’s two strategies — otherwise Predict would be quietly doing chain-of-thought too, and the comparison would be between reasoning and reasoning.
The fixture is the 37 cases of the training and development splits built in chapter 7. Each strategy ran three times over all 37, giving 111 measured calls per strategy and 222 in total. The holdout split was not touched.
That last constraint is not ceremony. A comparison run on cases you have already looked at tells you about those cases. Chapter 7 explains the discipline and chapter 18 demonstrates what happens without it.
The core of the harness is unremarkable:
def compare_modules(cases, programs, repetitions=3):
rows = []
for rep in range(repetitions):
for case in cases:
for name, program in programs.items():
pred = run_case(program, case)
rows.append({
"case_id": case["case_id"],
"module": name,
"rep": rep,
"rewritten_text": pred.rewritten_text,
})
return rows
The measured runner adds balanced execution order, latency and token accounting, output validation, provenance, and holdout protection. For each case and repetition it alternates which strategy runs first, so one strategy does not systematically receive the warmer position.
Balanced order reduces an obvious source of bias; it does not make the local model deterministic. Later repeated sessions show that call history can still move some outputs even at temperature zero. That is why small aggregate deltas need more caution than the raw precision of four decimal places suggests.
3. What we measured
| Measure | Predict | ChainOfThought |
|---|---|---|
| Measured calls | 111 | 111 |
| Hard-gate failures | 6 | 6 |
| Mean score, metric v1 | 0.8034 | 0.8175 |
| Mean score, metric v2 | 0.7862 | 0.8175 |
| Mean latency | 1.879 s | 3.113 s |
| Prompt tokens | 38,481 | 41,478 |
| Completion tokens | 8,490 | 14,481 |
| Total tokens | 46,971 | 55,959 |
The cost side is straightforward and roughly what you would expect. ChainOfThought took 66% longer, used 19% more total tokens, and 71% more completion tokens — it thinks out loud, and out loud costs money. Per case, the latency ratio ranged from 1.11× to 2.27×.
The quality side is where it gets interesting.
Under metric v1 — the structural/reference-overlap metric used for the Chapter 8 baseline — ChainOfThought scores 0.0141 higher. That clears the book’s old mechanical tie threshold of 0.01, but it is on the same scale as the cross-session movement later observed from the unchanged baseline.
So the important conclusion is not that +0.0141 is a stable quality gain. Taken alone, this aggregate is too small to carry that claim, and it does not justify 66% more latency. The more auditable evidence comes from asking which individual cases caused the two strategies and the two metrics to disagree.
Under metric v2 — which adds a semantic check on top of the same structural score — the gap is 0.0313, more than twice as large.
Note also that ChainOfThought’s v1 and v2 scores are identical to four decimal places. Across 111 calls it produced not one semantic violation. Predict’s two scores differ by 0.0172.
That difference is the whole finding, and it comes from one case.
4. The case that decided it
Case ed-025 belongs to the legal plain-language family. The task is to simplify a formal notice requirement without changing when its 30-day deadline begins.
The exact fixture sentence is:
Notification of the aforementioned defect must be made in writing not later than thirty (30) days following delivery.
The reference rewrite is:
You must report the defect in writing within 30 days of delivery.
Predict returned:
You must notify the sender in writing about the defect within 30 days of receiving the item.
ChainOfThought returned:
You must notify in writing about the defect within 30 days of delivery.
Both outputs perform the requested plain-language transformation. Both preserve the 30-day period and the requirement that notice be in writing. Under v1, Predict scored 0.933 and ChainOfThought scored 0.967.
The difference is the trigger event. The fixture explicitly requires that the deadline runs from delivery. Predict changed that to receiving the item. Those phrases can denote different points in time, so the rewrite violates a named semantic requirement even though it remains lexically close to the reference.
The semantic judge marked that requirement violated with reason code scope_changed, and metric v2 caps a rewrite with a violated semantic constraint at 0.30.
Now the arithmetic, which you can check:
ed-025 under v1: 0.933
ed-025 under v2: 0.300
difference: 0.633
spread across 37 cases: 0.633 / 37 = 0.0171
v2 delta − v1 delta: 0.0313 − 0.0141 = 0.0172
The entire v1-versus-v2 gap in this strategy comparison is driven by one distinct fixture case: ed-025. The case was run three times under each strategy, so this is not literally one observation, but it is one semantic situation out of 37.
That cuts both ways. The v2 result is unusually auditable: the aggregate difference reduces to a sentence pair, a frozen constraint, and arithmetic the reader can inspect. It is also fragile as a general claim about reasoning. Repeating the same case three times does not turn one kind of failure into a distribution of failure types.
So ed-025 is evidence that reasoning can avoid a semantic slip that direct prediction makes under this setup. It is not evidence that reasoning generally improves legal rewriting, semantic preservation, or one case in every 37.
5. Two metrics, two answers
Here is the same experiment, summarised twice.
Under v1: ChainOfThought scores 0.014 higher and costs 66% more time. The measured gain is too close to observed run-to-run movement to justify the additional cost on aggregate score alone.
Under v2: ChainOfThought scores 0.031 higher, produced no semantic violations in these 111 calls, and avoided the ed-025 trigger-date change that v1 rated highly.
Same 222 calls. Same outputs. Same latency measurements. What changes is the evaluation criterion. Under v1, the evidence does not justify paying for reasoning. Under v2, the experiment surfaces a specific correctness benefit whose value depends on the application’s cost of that failure.
That is not two contradictory datasets. It is one dataset answering two different evaluation questions.
This is the first time in this book that the choice of metric changes the answer rather than the confidence in the answer, and it will not be the last. In chapter 9 we attack v1 directly and find that three sentences engineered to reverse an author’s meaning score between 0.85 and 0.96. In chapters 11 and 12, two different optimizers independently discover that they can raise the v1 score by damaging a sentence, and both do it to the same case.
The pattern is already visible here in its mildest form: v1 rated a semantically invalid rewrite very highly. Predict received 0.933 after changing the deadline’s trigger event.
Nothing about the computation of 0.933 was a software error. The candidate changed the sentence, stayed within the scope-length rule, preserved the hard-gated terms, and overlapped strongly with the reference. v1 measured those properties correctly. It simply had no component capable of asking whether delivery had become receipt.
The metric answered the question encoded in its features. The engineering mistake would be to confuse that question with the whole task.
Hold on to that, because the temptation in chapters 10 through 13 will be to read a rising number as progress.
6. Reasoning traces are not evidence
A separate caution, and one the results above make easy to over-read.
ChainOfThought produces a reasoning field. It is tempting to treat that field as an explanation of why the answer is right. It is not. It is computational state produced by the same model, in the same call, as the answer — and a model that is about to make a mistake will happily produce fluent reasoning that leads to it.
For ed-002, the direct predictor returned:
Mira was tired from walking all day in rain and mud.
ChainOfThought reasoned explicitly that the rewrite should preserve the causal relationship while removing redundancy, and then returned the same sentence. The reasoning was correct and bought nothing.
The inverse case is more dangerous. A trace saying “the rewrite preserves all entities and improves rhythm” is not evidence that entities were preserved. It is a claim by an interested party. The only thing that established ed-025 as a failure was an evaluation run outside the program that produced it.
And even that is weaker than it sounds. The semantic judge in chapter 9 uses the same underlying model as the task program. It is a check, not an independent authority, and chapter 9 spends most of its length validating the judge rather than assuming it.
A richer execution policy is a hypothesis. A reasoning trace is not a result.
7. When to reach for reasoning
Diagnose the failure before paying for the strategy.
| Failure | Will reasoning help? | Better first move |
|---|---|---|
| Output field missing or malformed | Rarely | Strengthen field descriptions and validation |
| Entity renamed | Sometimes | Add an explicit constraint and a deterministic check |
| Meaning or scope quietly changed | Possibly; one case here improved with reasoning | Add a semantic check first, then measure whether reasoning reduces the failure |
| Sentence over-compressed | Sometimes | Supply more context first, then compare strategies |
| Type coercion failure | Rarely | Simplify the output or fix the parsing boundary |
| Wrong domain knowledge | No | Supply retrieval or change the model |
| Too slow | No | Prefer direct prediction or a smaller model |
The revised rule from this chapter’s measurement is narrower than “reasoning helps” and more useful. The aggregate v1 gain is small relative to observed run-to-run movement. The larger v2 difference is traceable to one distinct case in which reasoning preserved a semantic requirement that direct prediction changed.
That does not mean the other 36 cases were behaviorally identical, and it does not establish a general rate at which reasoning prevents semantic errors. It means we found one concrete failure class for which the additional computation mattered.
Whether that trade is worth paying depends on application stakes and failure cost. The experiment supplies the measured latency, token cost, and observed failure. Product or domain policy must decide how much that failure is worth avoiding.
Note also the six hard-gate failures, identical under both strategies. In this metric, the hard gate checks two deterministic properties: required entities must still be present, and forbidden terms must not appear.
Reasoning did not reduce those failures in this run. That does not prove reasoning can never help entity preservation or forbidden-term avoidance; it shows that changing only the execution policy left this measured failure count unchanged.
More importantly, these properties already have deterministic detectors. Once a failure can be checked directly, the system should not rely on a reasoning trace to certify it.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| Chain-of-thought looks better but outputs are longer | Style changed implicitly along with strategy | Compare outputs under the same length constraints | Put style constraints in the contract or the metric |
| The two strategies return different fields | The contract was not actually shared | Inspect both module definitions | Use the same signature class for both |
| Reasoning text reaches users | Intermediate state leaked into the response | Check API payloads and UI bindings | Return only the declared output fields |
| A tiny score difference is reported as a win | The delta is inside run-to-run noise | Repeat the run and measure the spread first | Establish a noise floor before quoting deltas |
| The model’s native reasoning is on | Model-level and framework-level reasoning are confounded | Check provider settings for thinking or reasoning modes | Disable it, or measure with it explicitly held constant |
| A rewrite scores well and is wrong | The metric cannot see the failure class | Inspect high-scoring outputs by hand, not just low-scoring ones | Add a check that can see it, then attack that check |
Conclusion
We can now run one task contract under different execution policies and attribute the difference to the policy. That is the mechanism this chapter set out to establish, and it worked.
What it produced was not the result we expected. Across 222 measured local calls, ChainOfThought cost 66% more time and 71% more completion tokens for a lexical improvement of 0.014 — near enough to noise that under the metric this book started with, the sensible conclusion was don’t bother.
Under a metric that can see meaning, the same 222 calls say something different, and the difference reduces to one sentence about a 30-day return window that Predict got subtly and expensively wrong.
Two lessons, and the second is the one that lasts.
The narrow one: extra reasoning is a cost you should price against specific failure modes, not against a generic belief that “reasoning is better.” Here the aggregate improvement was small, while one inspectable case showed a real semantic save that the original metric could not see.
The broader one: the measurement decided the answer. Nothing about the program changed between the two readings of this experiment. The only thing that changed was what we were able to see, and it flipped an engineering recommendation. Chapter 9 will attack the metric on purpose, and chapters 11 through 13 will hand it to optimizers whose entire job is to maximise it.
Before any of that, though, there is a more basic architectural question we have been deferring. Real systems rarely consist of one call. An editor might analyse the context, produce a candidate, then assess it — three stages, each with its own contract.
Decomposition is supposed to be obviously good. Chapter 5 measures it.
A production chain-of-thought module that carries reasoning style as an input rather than baking it into the program is excerpted in the Chapter 21 appendix.