← DSPy From First Principles

Separate What From How

Hold the task contract completely fixed, change only the execution strategy, and measure what the extra reasoning actually bought across 222 local calls.

Chapters 2 and 3 built a contract and said nothing about how a model should satisfy it. That was deliberate, and this chapter collects the payment.

Three objects must remain distinct throughout the comparison:

task contract       what valid behavior must preserve
execution strategy  how the model attempts the task
evaluation          how observed behavior is judged

The signature holds the first fixed. Predict and ChainOfThought vary the second. The frozen fixture and metrics supply the third. If two of those change together, the result no longer tells us what the reasoning policy caused.

                SAME CONTRACT

input ────────> Signature ────────> output
                    │
             execution policy
              /           \
         Predict       ChainOfThought

Because the task-facing contract is fixed, swapping the execution strategy gives us one declared experimental factor: Predict versus ChainOfThought. The strategy change will have consequences — a different rendered prompt, a reasoning field, different token use, different latency — but those are effects of the intervention rather than extra knobs changed by hand.

That is rare enough in language-model work to be worth stopping on. Many comparisons between prompting strategies change the instruction, the examples, the output format, and the strategy simultaneously, and then report which bundle felt better.

This chapter runs the comparison properly — 222 measured calls on a local model — and gets a result that surprised us, for a reason that turns out to be the subject of the rest of the book.


1. Two policies, one contract

Direct prediction:

import dspy

class RewriteSentence(dspy.Signature):
    """Rewrite one sentence to satisfy an editorial goal
    while preserving meaning, entities, and voice."""

    sentence: str = dspy.InputField(desc="The exact sentence to rewrite")
    goal: str = dspy.InputField(desc="The local reason this sentence is being edited")
    context: str = dspy.InputField(desc="Nearby prose needed to preserve continuity and voice")

    rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
    rationale: str = dspy.OutputField(desc="Short explanation of the edit")

direct = dspy.Predict(RewriteSentence)

Explicit reasoning, same signature:

reasoned = dspy.ChainOfThought(RewriteSentence)

ChainOfThought extends the signature with an additional reasoning output that the model fills before producing the declared fields. The contract does not change:

sentence, goal, context  →  rewritten_text, rationale

The policy does:

Predict:        answer directly
ChainOfThought: produce intermediate reasoning, then answer

So the calling code is identical for both, which is the property that makes the experiment possible:

def run_case(program, case):
    return program(
        sentence=case["sentence"],
        goal=case["goal"],
        context=case.get("context", ""),
    )

One implementation note. The shared repository RewriteSentence used by the measured runner already carries two additional inputs, issue_summary and preservation_notes, which Chapter 5 uses when the program is decomposed. The listing above shows the task-facing core for clarity. In this experiment both additional fields are passed as empty strings to both strategies.

So Predict and ChainOfThought receive the exact same expanded signature and the exact same values. The only deliberate intervention is the execution policy.


2. Designing a comparison that means something

The experiment is worth describing before the results, because the design is most of the value.

Held fixedVaried
Signature and field descriptionsThe module: Predict or ChainOfThought
Model, quantization, and generation settings
Cases and repetition count
Balanced execution-order policy
v1 and v2 metric definitions
Semantic-judge configuration for v2

The model was Qwen3 8.2B, Q4_K_M, served through Ollama and called via DSPy 3.3.1 at temperature 0. Qwen3’s native thinking mode was disabled, so that model-level reasoning would not contaminate the distinction between DSPy’s two strategies — otherwise Predict would be quietly doing chain-of-thought too, and the comparison would be between reasoning and reasoning.

The fixture is the 37 cases of the training and development splits built in chapter 7. Each strategy ran three times over all 37, giving 111 measured calls per strategy and 222 in total. The holdout split was not touched.

That last constraint is not ceremony. A comparison run on cases you have already looked at tells you about those cases. Chapter 7 explains the discipline and chapter 18 demonstrates what happens without it.

The core of the harness is unremarkable:

def compare_modules(cases, programs, repetitions=3):
    rows = []
    for rep in range(repetitions):
        for case in cases:
            for name, program in programs.items():
                pred = run_case(program, case)
                rows.append({
                    "case_id": case["case_id"],
                    "module": name,
                    "rep": rep,
                    "rewritten_text": pred.rewritten_text,
                })
    return rows

The measured runner adds balanced execution order, latency and token accounting, output validation, provenance, and holdout protection. For each case and repetition it alternates which strategy runs first, so one strategy does not systematically receive the warmer position.

Balanced order reduces an obvious source of bias; it does not make the local model deterministic. Later repeated sessions show that call history can still move some outputs even at temperature zero. That is why small aggregate deltas need more caution than the raw precision of four decimal places suggests.


3. What we measured

MeasurePredictChainOfThought
Measured calls111111
Hard-gate failures66
Mean score, metric v10.80340.8175
Mean score, metric v20.78620.8175
Mean latency1.879 s3.113 s
Prompt tokens38,48141,478
Completion tokens8,49014,481
Total tokens46,97155,959

The cost side is straightforward and roughly what you would expect. ChainOfThought took 66% longer, used 19% more total tokens, and 71% more completion tokens — it thinks out loud, and out loud costs money. Per case, the latency ratio ranged from 1.11× to 2.27×.

The quality side is where it gets interesting.

Under metric v1 — the structural/reference-overlap metric used for the Chapter 8 baseline — ChainOfThought scores 0.0141 higher. That clears the book’s old mechanical tie threshold of 0.01, but it is on the same scale as the cross-session movement later observed from the unchanged baseline.

So the important conclusion is not that +0.0141 is a stable quality gain. Taken alone, this aggregate is too small to carry that claim, and it does not justify 66% more latency. The more auditable evidence comes from asking which individual cases caused the two strategies and the two metrics to disagree.

Under metric v2 — which adds a semantic check on top of the same structural score — the gap is 0.0313, more than twice as large.

Note also that ChainOfThought’s v1 and v2 scores are identical to four decimal places. Across 111 calls it produced not one semantic violation. Predict’s two scores differ by 0.0172.

That difference is the whole finding, and it comes from one case.


4. The case that decided it

Case ed-025 belongs to the legal plain-language family. The task is to simplify a formal notice requirement without changing when its 30-day deadline begins.

The exact fixture sentence is:

Notification of the aforementioned defect must be made in writing not later than thirty (30) days following delivery.

The reference rewrite is:

You must report the defect in writing within 30 days of delivery.

Predict returned:

You must notify the sender in writing about the defect within 30 days of receiving the item.

ChainOfThought returned:

You must notify in writing about the defect within 30 days of delivery.

Both outputs perform the requested plain-language transformation. Both preserve the 30-day period and the requirement that notice be in writing. Under v1, Predict scored 0.933 and ChainOfThought scored 0.967.

The difference is the trigger event. The fixture explicitly requires that the deadline runs from delivery. Predict changed that to receiving the item. Those phrases can denote different points in time, so the rewrite violates a named semantic requirement even though it remains lexically close to the reference.

The semantic judge marked that requirement violated with reason code scope_changed, and metric v2 caps a rewrite with a violated semantic constraint at 0.30.

Now the arithmetic, which you can check:

ed-025 under v1:  0.933
ed-025 under v2:  0.300
difference:       0.633

spread across 37 cases:  0.633 / 37 = 0.0171

v2 delta − v1 delta:     0.0313 − 0.0141 = 0.0172

The entire v1-versus-v2 gap in this strategy comparison is driven by one distinct fixture case: ed-025. The case was run three times under each strategy, so this is not literally one observation, but it is one semantic situation out of 37.

That cuts both ways. The v2 result is unusually auditable: the aggregate difference reduces to a sentence pair, a frozen constraint, and arithmetic the reader can inspect. It is also fragile as a general claim about reasoning. Repeating the same case three times does not turn one kind of failure into a distribution of failure types.

So ed-025 is evidence that reasoning can avoid a semantic slip that direct prediction makes under this setup. It is not evidence that reasoning generally improves legal rewriting, semantic preservation, or one case in every 37.


5. Two metrics, two answers

Here is the same experiment, summarised twice.

Under v1: ChainOfThought scores 0.014 higher and costs 66% more time. The measured gain is too close to observed run-to-run movement to justify the additional cost on aggregate score alone.

Under v2: ChainOfThought scores 0.031 higher, produced no semantic violations in these 111 calls, and avoided the ed-025 trigger-date change that v1 rated highly.

Same 222 calls. Same outputs. Same latency measurements. What changes is the evaluation criterion. Under v1, the evidence does not justify paying for reasoning. Under v2, the experiment surfaces a specific correctness benefit whose value depends on the application’s cost of that failure.

That is not two contradictory datasets. It is one dataset answering two different evaluation questions.

This is the first time in this book that the choice of metric changes the answer rather than the confidence in the answer, and it will not be the last. In chapter 9 we attack v1 directly and find that three sentences engineered to reverse an author’s meaning score between 0.85 and 0.96. In chapters 11 and 12, two different optimizers independently discover that they can raise the v1 score by damaging a sentence, and both do it to the same case.

The pattern is already visible here in its mildest form: v1 rated a semantically invalid rewrite very highly. Predict received 0.933 after changing the deadline’s trigger event.

Nothing about the computation of 0.933 was a software error. The candidate changed the sentence, stayed within the scope-length rule, preserved the hard-gated terms, and overlapped strongly with the reference. v1 measured those properties correctly. It simply had no component capable of asking whether delivery had become receipt.

The metric answered the question encoded in its features. The engineering mistake would be to confuse that question with the whole task.

Hold on to that, because the temptation in chapters 10 through 13 will be to read a rising number as progress.


6. Reasoning traces are not evidence

A separate caution, and one the results above make easy to over-read.

ChainOfThought produces a reasoning field. It is tempting to treat that field as an explanation of why the answer is right. It is not. It is computational state produced by the same model, in the same call, as the answer — and a model that is about to make a mistake will happily produce fluent reasoning that leads to it.

For ed-002, the direct predictor returned:

Mira was tired from walking all day in rain and mud.

ChainOfThought reasoned explicitly that the rewrite should preserve the causal relationship while removing redundancy, and then returned the same sentence. The reasoning was correct and bought nothing.

The inverse case is more dangerous. A trace saying “the rewrite preserves all entities and improves rhythm” is not evidence that entities were preserved. It is a claim by an interested party. The only thing that established ed-025 as a failure was an evaluation run outside the program that produced it.

And even that is weaker than it sounds. The semantic judge in chapter 9 uses the same underlying model as the task program. It is a check, not an independent authority, and chapter 9 spends most of its length validating the judge rather than assuming it.

A richer execution policy is a hypothesis. A reasoning trace is not a result.


7. When to reach for reasoning

Diagnose the failure before paying for the strategy.

FailureWill reasoning help?Better first move
Output field missing or malformedRarelyStrengthen field descriptions and validation
Entity renamedSometimesAdd an explicit constraint and a deterministic check
Meaning or scope quietly changedPossibly; one case here improved with reasoningAdd a semantic check first, then measure whether reasoning reduces the failure
Sentence over-compressedSometimesSupply more context first, then compare strategies
Type coercion failureRarelySimplify the output or fix the parsing boundary
Wrong domain knowledgeNoSupply retrieval or change the model
Too slowNoPrefer direct prediction or a smaller model

The revised rule from this chapter’s measurement is narrower than “reasoning helps” and more useful. The aggregate v1 gain is small relative to observed run-to-run movement. The larger v2 difference is traceable to one distinct case in which reasoning preserved a semantic requirement that direct prediction changed.

That does not mean the other 36 cases were behaviorally identical, and it does not establish a general rate at which reasoning prevents semantic errors. It means we found one concrete failure class for which the additional computation mattered.

Whether that trade is worth paying depends on application stakes and failure cost. The experiment supplies the measured latency, token cost, and observed failure. Product or domain policy must decide how much that failure is worth avoiding.

Note also the six hard-gate failures, identical under both strategies. In this metric, the hard gate checks two deterministic properties: required entities must still be present, and forbidden terms must not appear.

Reasoning did not reduce those failures in this run. That does not prove reasoning can never help entity preservation or forbidden-term avoidance; it shows that changing only the execution policy left this measured failure count unchanged.

More importantly, these properties already have deterministic detectors. Once a failure can be checked directly, the system should not rely on a reasoning trace to certify it.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
Chain-of-thought looks better but outputs are longerStyle changed implicitly along with strategyCompare outputs under the same length constraintsPut style constraints in the contract or the metric
The two strategies return different fieldsThe contract was not actually sharedInspect both module definitionsUse the same signature class for both
Reasoning text reaches usersIntermediate state leaked into the responseCheck API payloads and UI bindingsReturn only the declared output fields
A tiny score difference is reported as a winThe delta is inside run-to-run noiseRepeat the run and measure the spread firstEstablish a noise floor before quoting deltas
The model’s native reasoning is onModel-level and framework-level reasoning are confoundedCheck provider settings for thinking or reasoning modesDisable it, or measure with it explicitly held constant
A rewrite scores well and is wrongThe metric cannot see the failure classInspect high-scoring outputs by hand, not just low-scoring onesAdd a check that can see it, then attack that check

Conclusion

We can now run one task contract under different execution policies and attribute the difference to the policy. That is the mechanism this chapter set out to establish, and it worked.

What it produced was not the result we expected. Across 222 measured local calls, ChainOfThought cost 66% more time and 71% more completion tokens for a lexical improvement of 0.014 — near enough to noise that under the metric this book started with, the sensible conclusion was don’t bother.

Under a metric that can see meaning, the same 222 calls say something different, and the difference reduces to one sentence about a 30-day return window that Predict got subtly and expensively wrong.

Two lessons, and the second is the one that lasts.

The narrow one: extra reasoning is a cost you should price against specific failure modes, not against a generic belief that “reasoning is better.” Here the aggregate improvement was small, while one inspectable case showed a real semantic save that the original metric could not see.

The broader one: the measurement decided the answer. Nothing about the program changed between the two readings of this experiment. The only thing that changed was what we were able to see, and it flipped an engineering recommendation. Chapter 9 will attack the metric on purpose, and chapters 11 through 13 will hand it to optimizers whose entire job is to maximise it.

Before any of that, though, there is a more basic architectural question we have been deferring. Real systems rarely consist of one call. An editor might analyse the context, produce a candidate, then assess it — three stages, each with its own contract.

Decomposition is supposed to be obviously good. Chapter 5 measures it.

A production chain-of-thought module that carries reasoning style as an input rather than baking it into the program is excerpted in the Chapter 21 appendix.