← DSPy From First Principles

Build Programs From Programs

Compose several DSPy modules with deterministic Python, then measure which intermediate stages actually influence the result and whether the decomposition earns its cost.

Chapter 4 held the contract fixed and changed the execution policy. This chapter changes the architecture.

Real applications rarely consist of one call. An editorial system might want to work out what is wrong with a sentence, produce a candidate, assess whether that candidate is risky, and hand the whole package to something that stores it.

    flowchart LR
    I[input] --> A[analyze]
    A --> W[rewrite]
    W --> S[assess]
    S --> R[return]
  

That is the pipeline as pitched: four stages in a line, each feeding the next. Decomposition is one of the few architectural moves that is widely treated as obviously correct. Smaller pieces, clearer boundaries, better debugging — the diagram sells itself.

We built it and measured it, and the results are worth the chapter. One intermediate stage has a measurable causal effect on the rewrite. Another has no causal path back to the rewrite at all.

The useful question therefore turns out not to be “does decomposition help?” It is “which edge in the graph changes the outcome, which stage merely observes it, and what does each one cost?”


1. The giant prompt version

The alternative to decomposition is to ask for everything at once:

Analyze this sentence, identify issues, rewrite it, assess risk,
explain your decision, and return JSON...

This sometimes works. It is hard to debug, because when the final rewrite is bad you cannot tell which part of that instruction failed:

context interpretation?
issue diagnosis?
rewrite generation?
risk assessment?
output formatting?

A single prompt compresses several operations into one opaque step. Decomposition gives you places to look. It also adds calls, latency, schemas, and integration surface, so the question is not whether decomposition is cleaner — it is whether the intermediate state is worth what it costs.

That question has a measurable answer, which is the point of this chapter.


2. Three small contracts

import dspy

class AnalyzeSentence(dspy.Signature):
    """Identify edit issues and preservation notes for a sentence rewrite."""

    sentence: str = dspy.InputField(desc="One sentence to improve.")
    goal: str = dspy.InputField(desc="Editorial goal for the rewrite.")
    context: str = dspy.InputField(desc="Decision-time local context.")

    issue_summary: str = dspy.OutputField(desc="Concise explanation of the edit issue.")
    preservation_notes: str = dspy.OutputField(desc="Decision-time details that should be preserved.")

class RewriteSentence(dspy.Signature):
    """Rewrite a sentence using the analysis and preservation notes."""

    sentence: str = dspy.InputField(desc="Original sentence.")
    goal: str = dspy.InputField(desc="Editorial goal for the rewrite.")
    context: str = dspy.InputField(desc="Decision-time local context.")
    issue_summary: str = dspy.InputField(desc="Analysis of the edit issue.")
    preservation_notes: str = dspy.InputField(desc="Details to preserve.")

    rewritten_text: str = dspy.OutputField(desc="One rewritten sentence.")
    rationale: str = dspy.OutputField(desc="Brief reason for the rewrite.")

class AssessRewrite(dspy.Signature):
    """Assess risk in a proposed rewrite."""

    sentence: str = dspy.InputField(desc="Original sentence.")
    rewritten_text: str = dspy.InputField(desc="Candidate rewritten sentence.")
    goal: str = dspy.InputField(desc="Editorial goal.")
    context: str = dspy.InputField(desc="Decision-time local context.")

    risk: str = dspy.OutputField(desc="Risk level: low, medium, or high.")
    assessment: str = dspy.OutputField(desc="Brief explanation of preservation risk.")

RewriteSentence is the signature from chapters 2 through 4, now carrying the two extra inputs the analysis stage provides. Chapter 4 passed those as empty strings, which is exactly the BLANK condition we will use as a baseline in section 6.

These are three program boundaries, not three prompts. Each has its own inputs, its own outputs, and its own place to fail.


3. Compose them in a custom module

class EditorialRewriteProgram(dspy.Module):
    def __init__(self) -> None:
        super().__init__()
        self.analyze = dspy.Predict(AnalyzeSentence)
        self.rewrite = dspy.Predict(RewriteSentence)
        self.assess = dspy.Predict(AssessRewrite)

    def forward(self, sentence: str, goal: str, context: str = "") -> dspy.Prediction:
        analysis = self.analyze(sentence=sentence, goal=goal, context=context)

        rewrite = self.rewrite(
            sentence=sentence,
            goal=goal,
            context=context,
            issue_summary=analysis.issue_summary,
            preservation_notes=analysis.preservation_notes,
        )

        assessment = self.assess(
            sentence=sentence,
            rewritten_text=rewrite.rewritten_text,
            goal=goal,
            context=context,
        )

        return dspy.Prediction(
            rewritten_text=rewrite.rewritten_text,
            rationale=rewrite.rationale,
            risk=assessment.risk,
            assessment=assessment.assessment,
            issue_summary=analysis.issue_summary,
            preservation_notes=analysis.preservation_notes,
        )

This is the same EditorialRewriteProgram from chapter 2, and its external contract has not changed. Callers still pass sentence, goal, context and still read rewritten_text. The internals went from one LM call to three and nothing upstream noticed — which was the promise chapter 2 made about the module boundary, now delivered.

Look closely at the data flow, because it matters in section 7.

    flowchart LR
    A[analyze] -->|issue_summary, preservation_notes| W[rewrite]
    W -->|rewritten_text| R[returned sentence]
    W --> S[assess]
    S -.->|risk and assessment, returned but unused| D[no retry, rejection, or branch]
  

The analysis feeds the rewrite. The rewrite feeds both the assessment and the returned sentence. The assessment feeds nothing that can change the output: it produces risk and assessment, the program returns them, and no code path uses either value to alter the sentence that gets returned.


4. Not everything should be an LM call

A DSPy module holds ordinary Python as happily as it holds predictors, and some questions should never reach a model:

def has_same_terminal_punctuation(original: str, candidate: str) -> bool:
    return original.rstrip()[-1:] == candidate.rstrip()[-1:]

def too_much_change(original: str, candidate: str, limit: float = 0.6) -> bool:
    before = original.split()
    after = candidate.split()
    delta = abs(len(after) - len(before))
    return delta / max(len(before), 1) > limit

These are crude and they illustrate the boundary. If a question is computable, compute it. Do not spend a network round trip and 400 tokens asking a language model whether two strings end with the same character, and do not trust its answer when you do.

LM:
  proposes language

Python:
  routes, validates, stores, hashes, compares, enforces policy

This division holds all the way up. In Chapter 8 the hard gate is pure Python: required entities must remain present and forbidden terms must remain absent. Other output-shape and scope checks are recorded separately rather than being smuggled into that gate.

In Chapter 9 the semantic component uses an LM judge because a requirement such as “the deadline still runs from delivery” is not reducible to token overlap. In Chapter 19 the promotion decision is deterministic code reading measured evidence, because the program that proposes a candidate should not also be the authority that decides whether the candidate ships.

The rule that generalises: the model proposes, the software decides.


5. What the composed program measured

We ran EditorialRewriteProgram across the 37 training and development cases. Each execution makes three LM calls.

MeasureFresh composed runChapter 4 direct predictor
LM calls per case31
Mean score, metric v10.83590.8034
Mean score, metric v20.83590.7862
Semantic violations01

The fresh composed run sits about 0.033 above the Chapter 4 direct-predictor point estimate on v1. Under v2 the point-estimate gap is larger because the direct predictor’s ed-025 semantic violation is capped by v2.

Do not turn those two rows into a precise causal effect yet. They come from different sessions, and the later repeat work shows that local Qwen3 can move between session- and order-dependent attractors even at temperature zero.

The earlier composition run recorded about 8.0 seconds per case, versus 1.879 seconds for Chapter 4’s direct predictor, which suggests a cost on the order of four times the latency. The final ratio should be quoted from the paired repeated timing run rather than by combining timing from one session with quality from another.

Before drawing conclusions from that, an admission that is more useful than the numbers.

We ran the same composed program in separate sessions and got 0.7794 and 0.8359.

That 0.0565 spread is large enough to reverse the architectural story if you treat either run as a definitive point estimate. Calling it a simple noise floor would be misleading. The later repeat work finds both cross-session drift and order-dependent drift under interleaving, even at temperature zero.

Here the aggregate movement is easy to localise. The low-scoring run hit two hard-gate zeros — ed-010 and ed-020 — that the higher-scoring path did not. Because a hard gate maps a candidate straight to zero, a small change in generated text can become a large change in the mean.

The arithmetic is unforgiving. Replacing two otherwise roughly 0.83 cases with hard zeros moves a 37-case mean by about 2 × 0.83 / 37 ≈ 0.045. One borderline case can therefore move the aggregate by roughly 0.022 — larger than some of the optimizer deltas we later spend whole chapters interpreting.

Three rules follow.

Report gate failures beside the mean. A score without its failure count hides whether the aggregate moved because many cases improved slightly or one case crossed a discontinuous boundary.

Repeat comparisons that depend on hard-zero cases. We nearly treated 0.7794 as the chapter’s headline and would have concluded that composition was harmful. A second session changed that story.

Do not equate temperature=0 with determinism. In the run comparison, 27 of 37 generated analysis pairs changed across sessions. On ed-010, even byte-identical analysis input led to different rewrites: one preserved forbidden present-tense forms and hard-failed; another converted the sentence to past tense and scored about 0.967.

The instability is therefore not confined to one stage or one kind of state. Experimental design has to measure it rather than assume it away.


6. Data flow is not causal influence

The composed program produces intermediate state that can be audited. A lineage check confirms that analysis.issue_summary reached the rewrite unchanged, that analysis.preservation_notes did too, and that rewrite.rewritten_text reached the assessor unchanged. All lineage checks pass.

That proves the plumbing. It says nothing about whether the analysis is doing any work.

To find out, we held the sentence, goal, context, model, signature, and generation settings fixed and varied only the two analysis fields entering the rewrite stage:

Conditionissue_summarypreservation_notes
REALAnalysis from this caseAnalysis from this case
BLANKEmptyEmpty
SHUFFLEDAnalysis from a different caseAnalysis from a different case

The ablation evaluates 37 cases across three conditions, for 111 measured rewrite calls. The REAL analysis was generated separately and then supplied to the rewrite stage, so the latency column below prices only the rewrite call. It does not include the cost of producing the analysis that REAL depends on.

That distinction matters when we later ask whether the analysis stage earns its call: the quality effect and the end-to-end cost come from different measurements and should not be collapsed into one number.

ConditionScore v1Score v2Rewrite latencySemantic violations
REAL0.83540.83541.957 s0
BLANK0.80050.78341.819 s1
SHUFFLED0.78450.77961.915 s1

In this 37-case ablation, REAL analysis scores about 0.035 above BLANK and 0.051 above SHUFFLED on v1. It is also the only condition in this run with zero semantic violations.

That is evidence that the intermediate analysis is causally useful under this fixture: changing only the analysis inputs changes measured rewrite quality in the expected direction.

It is not yet a universal claim that an analysis call always earns its cost. The ablation measures the benefit of having the correct analysis available; end-to-end deployment still has to price the extra Analyze call and repeat the comparison across sessions.


7. The stage that pays for nothing

Now put two numbers side by side.

A fresh full-pipeline run — Analyze, Rewrite, Assess — scores 0.8359. The ablation’s REAL path — Analyze, Rewrite, with no Assess call — scores 0.8354.

The near-equality is reassuring, but it is not the proof. These are separate runs with different call schedules, and the chapter has just shown why cross-run equality or inequality can be noisy.

The stronger evidence is the dataflow itself. assess runs only after rewrite.rewritten_text has already been produced. It receives that finished sentence and emits risk and assessment. Neither field feeds a retry, rejection, rewrite, or branch before the returned rewritten_text is fixed.

So under the program as written, the assessment stage has no causal path to the scored rewrite. Even if two noisy sessions had produced different aggregate means, that difference could not be attributed to assessment improving the sentence; the code gives it no mechanism to do so.

What it definitely has is a price: one additional LM call per case, plus its prompt and completion tokens. The available measurements establish substantial end-to-end overhead, but they do not yet justify claiming that assess is the single largest latency component without a stage-level timing breakdown.

Chapter 3 gave a rule: every field needs a job, and any field that can control an action needs independent validation or calibration. It used risk as the example of a field not to add merely because it sounds useful.

Here that rule is violated at the scale of a whole stage. A dedicated LM call produces risk and assessment; the program returns them; nothing in the measured path consumes them.

The outputs may still have diagnostic value. A human can inspect what the assessor claimed was risky, just as a human can inspect a rationale or reasoning trace. But that is observability, not quality control. Unless the assessment changes routing or a candidate decision, it should be priced as diagnostic instrumentation rather than credited with improving the rewrite.

Giving the stage a consumer would at least make it causal:

assessment = self.assess(...)

# Illustrative only: an unvalidated self-report should not be a
# production acceptance gate without calibration or another check.
if assessment.risk == "high":
    return dspy.Prediction(
        rewritten_text=sentence,          # decline the edit
        rationale=f"declined: {assessment.assessment}",
        risk=assessment.risk,
    )

Now risk can change the returned sentence, and returning the original is exactly the kind of restraint Chapters 2 and 3 argued the program should be able to represent explicitly.

But causal is not the same as safe. This branch would also give an unvalidated LM self-assessment authority over the output — the same design mistake Chapter 3 warned about with confidence. Before such a branch controlled acceptance, risk would need calibration, an independent validator, or a safer policy such as routing high-risk cases to human review rather than automatically deciding them.

We have not measured any of those variants, so they remain design sketches rather than results.

The measured conclusion is narrower and stronger:

A stage with no downstream consumer cannot change the scored rewrite. It can add observability, state, and cost — but not control.

For now, we make the structural decision ourselves: keep the useful analysis path and recognise that the assessor has no causal authority. Later the book will ask a harder question — whether program structure itself can become an optimization variable, and whether an optimizer can discover that a stage is decorative rather than merely tuning the words inside it.


8. Four composition patterns

Where should one module end and the next begin?

Linear:  A -> B -> C
Fan-out: A -> B, C, D
Wrapper: Memory(ContextPolicy(Module))
Graph:   explicit dependency topology

Linear pipelines suit work where each stage produces state the next genuinely needs — which, as section 6 shows, is a claim to test rather than assume. Fan-out suits independent specialists working over the same evidence. Wrappers suit cross-cutting concerns: context policy, memory, provenance, retries, cost tracking. Graphs suit real dependencies, parallel waves, checkpoints, and recovery.

Three public DSPy systems are worth reading for these shapes. CodeSpy shows wrapper and specialist composition in a code-review system, including memory and context wrappers around parallel review roles. STORM separates knowledge curation, outlining, article generation, and polishing into distinct modules. Arachne is the graph-shaped contrast: clarify a goal, weave a typed dependency graph, provision tools, execute in waves, evaluate, then heal or reroute.

Read them for the boundaries, and ask of each stage the question section 7 asks: what consumes this?


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
The composed program is slower than expectedA stage may produce state that never affects the decisionTrace each intermediate field to its downstream consumerRemove it, keep it explicitly as observability, or give it a validated role
Intermediate state flows correctly but changes nothingThe field is too vague to be actionableAblate it: REAL vs BLANK vs SHUFFLEDNarrow the field or remove the stage
Wrong intermediate state scores wellThe metric rewards a proxy, not the propertyScore the same outputs under a semantic checkStrengthen the metric before trusting the result
The mean moves sharply between temperature-zero runsSession/order effects push borderline cases across hard gatesCompare per-case outputs and gate failures, not only aggregatesRepeat in fresh sessions with balanced order; report gate counts
A three-case experiment gave a confident answerThe fixture cannot resolve the effect sizeCompute what delta the fixture could detectGet more cases before drawing the conclusion
Every step is an LM callComputable work was delegated to the modelSearch for calls that ask deterministic questionsMove them into Python
Debug payloads are enormousThe program stores raw everythingInspect a stored run recordStore bounded traces, fingerprints, and key fields

Conclusion

We have a real language-model program now: small modules, deterministic Python between them, intermediate state that can be inspected, traced, and perturbed.

We also have a more honest account of what decomposition bought. A fresh composed run scores about 0.836, while the Chapter 4 direct-predictor point estimate is about 0.803 on v1. Because those values come from different sessions, the exact end-to-end delta should be treated as provisional until the paired repeated comparison is complete.

The stage-level result is clearer. In the 37-case ablation, REAL analysis scores about 0.035 above BLANK and 0.051 above SHUFFLED on v1, with zero semantic violations in that run. The assessment stage, by contrast, has no causal path to rewritten_text because nothing consumes its outputs before the result is returned.

That is the useful shape of the answer. “Does decomposition help?” is the wrong question. Ask instead: Which stage changes behavior? Which stage changes only observability? What does each one cost? And which outputs are trusted enough to control anything?

Two assumptions fail in this chapter.

The first is that explicit intermediate state is valuable merely because it is explicit. The assessment stage is explicit, inspectable, and well named, yet under the measured program it cannot influence the scored rewrite at all. Its value, if any, is diagnostic unless we deliberately give it a validated downstream role.

Everything so far has treated the model as background — a dependency configured once and then ignored. Every number in chapters 4 and 5 is a number about Qwen3 8.2B at temperature zero, and none of them would necessarily survive a different model.

So:

Is the model part of the program, or a dependency of it?