A Prompt Is Not Yet a Program
Define the properties a language-model component needs before it deserves to be called a program — and see why those boundaries are prerequisites for controlled comparison.
Chapter 1 ended with a signature and an uncomfortable amount of unfinished business. We had established that one prompt string was carrying the task semantics, the behavioral framing, the output schema, and the provider assumptions simultaneously, and that this is why three unrelated failures arrived looking identical.
The obvious response is to write a better string. That is the wrong move, and it is worth being precise about why.
A prompt is text:
"Rewrite this sentence..."
A plain handwritten prompt has no explicit semantic seams. You can diff characters, split the string into templates, or give pieces different filenames, but unless the application declares what those pieces mean, the boundaries are only conventions. A change to the wording may alter the task, the execution tactic, the output format, or several of them at once.
A program has a boundary:
inputs
↓
declared behavior
↓
execution strategy
↓
structured outputs
The boundary is the entire point. It makes controlled change possible: task contract, execution strategy, model configuration, examples, and evaluation can become separately named variables.
That is the step from structure to experiment. Once the variables have names, we can hold four fixed, change one, and compare candidate A with candidate B without pretending that a diff between two prompt strings explains the cause.
flowchart LR
subgraph Prompt
PS[prompt string] --> B[one inseparable bundle]
end
subgraph Program
PB[program boundary] --> C[contract + strategy + model + data + evaluation]
C --> CC[controlled candidates]
end
A prompt string fuses every decision into one object. A program boundary keeps those decisions as separate parts, which is what makes a candidate controlled — you can change one part and hold the rest fixed. That separation does not by itself prove that only one thing changed. You still need fingerprints, run records, and a frozen evaluation protocol to establish that. But without the boundary, there is not even a stable place to ask the question.
This chapter is about what has to be true before a language-model component earns that description. It is deliberately not about making the signature good — that is chapter 3, and it turns out to be harder. This chapter is about the shape.
1. What a program promises
An ordinary Python function has a weak contract that still does real work:
def slugify(title: str) -> str:
...
That line says almost nothing about behavior. It says enough. It names the inputs, names the output type, gives other code something stable to call, creates a place where a test can live, and lets you replace the implementation tomorrow without telling anyone.
A language-model component becomes engineerable when the same concerns are made explicit:
| Property | Why it matters |
|---|---|
| Named inputs | Callers stop assembling task inputs as ad hoc prose |
| Declared behavior | The semantic job is visible outside prompt wording |
| Execution strategy | You know how the declared task is being attempted, separately from what it is |
| Structured outputs | Downstream code gets stable fields instead of a paragraph to parse |
| Inspectable structure | You can enumerate the program’s parts rather than infer them from a string |
| A place to fail | Validation, retries and fallbacks have somewhere to live |
DSPy provides two handles for this.
A signature declares the task interface: what comes in, what goes out, and the semantic job connecting them.
A module controls how that task is attempted — direct prediction, explicit reasoning, tool use, retrieval, or a composition of several calls with ordinary Python between them.
Signature: what behavior is requested?
Module: how is that behavior attempted?
That sentence is easy to nod at and easy to underestimate. The next section makes it concrete, because it is the idea the rest of the book is built on.
2. The separation that does the work
Here is the smallest program-shaped version of the running task. One sentence goes in and one rewrite comes out.
This is the pedagogical one-call form. In Chapter 5 the repository implementation expands the rewrite stage with intermediate analysis fields, but the core task boundary remains sentence, goal, and context in, with a rewrite and rationale out. Keeping that distinction explicit prevents the teaching interface from being mistaken for a claim that the implementation never evolves.
import dspy
class RewriteSentence(dspy.Signature):
"""Rewrite one sentence to satisfy an editorial goal
while preserving meaning, entities, and voice."""
sentence: str = dspy.InputField(desc="The exact sentence to rewrite")
goal: str = dspy.InputField(desc="The local reason this sentence is being edited")
context: str = dspy.InputField(desc="Nearby prose needed to preserve continuity and voice")
rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
rationale: str = dspy.OutputField(desc="Short explanation of the edit")
class EditorialRewriteProgram(dspy.Module):
def __init__(self) -> None:
super().__init__()
self.rewrite = dspy.Predict(RewriteSentence)
def forward(self, sentence: str, goal: str, context: str = "") -> dspy.Prediction:
return self.rewrite(sentence=sentence, goal=goal, context=context)
Calling it is unremarkable:
program = EditorialRewriteProgram()
result = program(
sentence="Jalen opened the door and then he looked into the room and felt afraid.",
goal="Sharpen the sentence without changing the event.",
context="A tense scene in a realist novel.",
)
print(result.rewritten_text)
Now watch what the separation buys. The same signature, attempted a different way:
direct = dspy.Predict(RewriteSentence)
reasoned = dspy.ChainOfThought(RewriteSentence)
Both take sentence, goal and context. Both return rewritten_text and rationale. ChainOfThought adds an internal reasoning step and constructs a substantially different prompt to do it. The task contract did not move, so the calling code does not move either:
for module in (direct, reasoned):
out = module(sentence=..., goal=..., context=...)
print(out.rewritten_text)
This is not a stylistic nicety. It is the thing that makes the question “is chain-of-thought worth it here?” answerable at all. When reasoning strategy is a phrase inside a prompt string, swapping it means rewriting the prompt, which means changing several things at once, which means the comparison is uninterpretable — exactly the v6-versus-v7 problem from chapter 1. When the strategy is a module wrapping a fixed signature, you can hold everything else still and measure the difference.
We do measure it in Chapter 4 on 37 train-and-development cases with three repetitions per strategy. ChainOfThought costs more, but the larger fixture no longer produces a clean tie: it has a small advantage under the original metric and a clearer advantage under the semantic guardrail because plain Predict changes the reference point of a legal deadline on one case.
The lesson is not that reasoning always wins. It is that separating the strategy from the contract makes a question like cost-versus-correctness measurable instead of rhetorical.
The same holds for structure. EditorialRewriteProgram starts here as one LM call. In Chapter 5 it becomes three — an analysis stage, a rewrite stage, and an assessment stage, with intermediate state passed between them. Its internals change completely. The code that calls the program can keep the same task-facing boundary because sentence, goal, and context still identify the work to be done.
The Chapter 5 experiments then separate two questions that are easy to confuse. Real intermediate analysis is causally useful: on the enlarged fixture it beats blank or shuffled analysis. But a useful intermediate representation does not automatically justify the cost of the whole three-call pipeline. Internal state can help one stage while the end-to-end program still has to earn its extra latency and tokens.
3. A program is several artifacts, not one
The handwritten version had one artifact:
prompt string
At this stage, the program gives us at least six named artifacts to reason about separately:
signature the task contract (chapters 2–3)
module the execution strategy (chapters 4–6)
LM configuration the model and parameters (chapter 6)
demonstrations examples, once compiled (chapters 10–13)
metric the acceptance criterion (chapters 8–9)
run record what actually happened (chapter 8)
These artifacts can be identified, fingerprinted, logged, and compared separately. They are not independent in the statistical sense: changing a signature can alter the prompts produced by a module, changing demonstrations can alter token cost, and changing the LM can alter every observed output. The gain is that those dependencies are visible rather than fused into one string.
Concretely, the handles you now have:
If downstream code needs a risk assessment as well as a rewrite, you change the contract and every caller sees the new field. If the model should reason before answering, you change the module while preserving the task-facing contract. If you move from one model to another, you change the LM configuration and record that as an experimental variable. If you later compile the program against examples, the optimizer has structured program state to modify rather than an undifferentiated prompt blob.
So when a number changes, you can ask which artifacts moved and which were held fixed. Later chapters add dataset, split, evaluation-protocol, candidate, promotion, and deployment identities around these six. That is not bookkeeping ceremony — it is what turns attribution from a guess into an experiment.
4. Declared behavior is not implementation
The most common failure when people first write signatures is to pour the old prompt into the docstring.
class BadRewriteSentence(dspy.Signature):
"""Think step by step, compare three alternatives, choose the best one,
and rewrite the sentence using a terse Hemingway-like style."""
sentence: str = dspy.InputField()
rewritten_text: str = dspy.OutputField()
That docstring mixes three different kinds of statement, and they belong in three different places:
| Statement | Kind | Where it belongs |
|---|---|---|
| “rewrite the sentence” | Required behavior | The signature. This is the task. |
| “think step by step, compare three alternatives” | Execution tactic | The module. This is a ChainOfThought decision, not a contract term. |
| “using a terse Hemingway-like style” | Editorial policy | The goal field or the data. This is one theory of good prose, not a property of the task. |
Fusing them has real costs. You cannot swap Predict for ChainOfThought and get a clean comparison, because the docstring is already asking for reasoning — you would be measuring reasoning against reasoning. You cannot use the program on an author who does not write like Hemingway. And when an optimizer later rewrites this instruction, as MIPROv2 does in chapter 12, it will be free to modify your task definition and your tactic and your style policy all at once, which puts you straight back in the position chapter 1 diagnosed.
A useful test: read the docstring and ask whether each clause would still be true if you swapped the execution strategy. “Rewrite the sentence to satisfy an editorial goal” survives that swap. “Think step by step” does not.
5. Outputs are part of the program
Output fields are not decoration. They determine what the rest of the system is able to do, and they determine what the model is able to say.
Start with the coarse version:
class CoarseRewrite(dspy.Signature):
"""Improve the sentence."""
sentence: str = dspy.InputField()
answer: str = dspy.OutputField()
A model given this will typically return something like:
Jalen opened the door, afraid of what he would find. I tightened the rhythm by
removing the coordinating conjunction and compressing the second clause, while
keeping the fear explicit.
Now the downstream code has to separate the rewrite from the commentary, which means a regex, which means a regex that works until the model puts the explanation first. Splitting the fields removes the problem rather than handling it:
rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
rationale: str = dspy.OutputField(desc="Short explanation of the edit")
That much is standard advice. Here is the version of it that is easy to miss, and that this book paid for.
A contract determines which distinctions the model can express cleanly — and which distinctions the rest of the system can observe.
Look again at the output field:
rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
The model can technically return the original sentence unchanged, because rewritten_text is still just a string. What the contract cannot express is the semantic distinction between “I deliberately decided no edit was needed” and “I attempted a rewrite and happened to reproduce the input.”
The field name and task wording also create pressure toward change. If restraint is sometimes the correct editorial decision, that state should be representable and evaluable deliberately rather than inferred from string equality after the fact.
This is not a hypothetical. In the expanded evaluation fixture, six of forty-four cases — about 13.6% — target unnecessary editing or restraint. On those cases, an unchanged or near-unchanged sentence can be the correct behavior.
The original structural metric used for the Chapter 8 baseline gives an unchanged output only 0.65, because its changed component is zero even when restraint is correct. Chapter 9 then makes the deeper problem visible: a metric can faithfully implement the assumptions built into the contract and still reward the wrong behavior.
The lesson is not that the current signature makes identity output impossible. It is that the program has no explicit representation of the decision not to edit, and the metric inherits that ambiguity. A weak contract can therefore propagate directly into a weak measurement system.
The fixes are ordinary once you see the problem:
class TriagedRewrite(dspy.Signature):
"""Decide whether a sentence needs editing, and rewrite it only if it does."""
sentence: str = dspy.InputField()
goal: str = dspy.InputField()
context: str = dspy.InputField()
edit_needed: bool = dspy.OutputField(desc="False if the sentence already satisfies the goal")
rewritten_text: str = dspy.OutputField(desc="The replacement sentence, or the original unchanged")
rationale: str = dspy.OutputField()
Or make rewritten_text == sentence an explicitly valid and expected outcome, documented in the field description and handled deliberately by the metric. Or split triage into its own stage.
Which of those is right depends on the application. The point is that all three are decisions about the contract, made before any model runs, and getting them wrong is not merely a prompt-wording problem — it is a shape problem that every downstream stage inherits.
For the experiments in this book, we deliberately keep the simpler rewrite contract long enough to expose that propagation. We do not retroactively repair the fixture after discovering the 0.65 restraint floor, because doing so would erase the evidence. The triaged signature above is a design alternative; the measured corpus remains a record of what the original contract got wrong.
6. The boundary is where failure and cost live
One more property, easy to overlook because a prompt has no equivalent.
A prompt cannot handle its own failure. If the model returns nonsense, the string has no opinion about that; whatever code called it has to cope, usually far from the point of the mistake. Chapter 1’s parse_json_object is a small monument to this — a function that exists in the caller because the prompt could not enforce anything.
A module can absorb failure at the point it occurs:
class EditorialRewriteProgram(dspy.Module):
def __init__(self) -> None:
super().__init__()
self.rewrite = dspy.Predict(RewriteSentence)
def forward(self, sentence: str, goal: str, context: str = "") -> dspy.Prediction:
result = self.rewrite(sentence=sentence, goal=goal, context=context)
if not result.rewritten_text.strip():
return dspy.Prediction(rewritten_text=sentence, rationale="empty output; original retained")
return result
That is a small thing and it demonstrates a structural one: there is now a place where “what should happen when this goes wrong” can be written down, versioned with the program, and visible to whoever reads it next.
The example handles only one deterministic failure: an empty rewrite. It does not detect an entity change, a semantic drift, or a stylistically wrong but fluent sentence. A module boundary creates a place for failure policy; it does not make failure detection automatic. Later chapters add explicit gates, metrics, judges, and promotion rules at that boundary.
The module boundary is also where cost accumulates, and this is worth naming early. Every decision about how has a price. ChainOfThought generates more tokens than Predict. A three-stage composition makes three calls instead of one. A retrieval step adds a lookup and a much longer prompt. An agent loop may make five calls or fifteen and you will not know which until it runs.
None of those costs is visible in the signature, which is exactly right — the contract should not know how it is being satisfied. But it means that from here on, every architectural choice in this book has to justify itself against a measured baseline rather than an intuition. Chapter 4 prices reasoning. Chapter 5 prices decomposition. Chapter 15 prices tool use. Chapter 16 prices retrieval, and finds one policy that costs ten times the tokens and scores worse.
The separation between what and how is what makes those prices legible. Without it, the cost of an architectural decision is just part of the bill.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| The signature docstring reads like a prompt | Task contract and execution tactic are fused | Check whether each clause survives swapping the module | Move tactic into the module, style into the data |
| Downstream code runs regexes on the output | The output field is too coarse | Search for parsing that happens after the LM call returns | Split the output into named fields |
| Every caller passes slightly different context prose | Inputs are unnamed or under-specified | Compare call sites for the same program | Name the input fields and document their scope |
| Swapping modules breaks callers | Tactic-specific details leaked into the signature | Look for fields named after reasoning steps, prompts or providers | Keep signature fields semantic |
| The model edits text that was already correct | The contract cannot express “no change needed” | Check whether the identity output is representable and expected | Add a triage output, or make unchanged text an explicit valid result |
| A regression appears and nobody can locate it | Several of the six artifacts changed together | Diff signature, module, LM config, demos and metric independently | Version them separately and change one at a time |
Conclusion
We now have a program-shaped boundary around a language-model behavior. The prompt still exists — DSPy builds one on every call, and you can print it — but it is no longer the thing the rest of the software has to understand or the thing you edit to change behavior.
More importantly, behavior can now become an experimental variable. A baseline and candidate can share a task contract while differing in execution strategy, demonstrations, or instructions. That comparison was not meaningful while every concern lived inside one string.
The assumption we removed is that a single string should carry the task contract, the execution tactic, the output interface, and the version identity as one indivisible artifact. Those concerns now have separate handles, which lets later chapters vary them under recorded protocols instead of relying on prompt-version folklore.
But notice what Section 5 exposed. We designed an output field in the most natural way available and failed to represent “no edit needed” as an explicit decision. Six of forty-four later fixture cases make restraint central, and the first structural metric penalises unchanged output even when restraint is correct. Nothing about having a signature prevented that.
The boundary is in the right place. The next question is whether the contract at that boundary expresses the right distinctions.
Which is the next problem:
What exactly should the model be asked to consume and produce — and how do you tell a strong contract from a weak one before it costs you five chapters?