<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>DSPy From First Principles on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/books/dspy-from-first-principles/</link><description>Recent content in DSPy From First Principles on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Thu, 03 Sep 2026 11:00:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/books/dspy-from-first-principles/index.xml" rel="self" type="application/rss+xml"/><item><title>Why Are We Still Hand-Writing Prompts?</title><link>https://aibussin.com/books/dspy-from-first-principles/01-chapter/</link><pubDate>Fri, 28 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/01-chapter/</guid><description>&lt;p&gt;Almost every language-model application starts as one string.&lt;/p&gt;&#10;&lt;p&gt;That is not a mistake. A prompt is the fastest way to find out whether a model can do the job at all, and it lets you work directly on the behavior instead of building a framework around a problem you do not yet understand. Most good LM systems begin this way and should.&lt;/p&gt;&#10;&lt;p&gt;The trouble is that the prompt usually survives longer than its usefulness. It stops being a probe and becomes the specification, and by the time anyone notices, the string is four hundred words long, nobody remembers why the third paragraph is there, and changing it feels dangerous.&lt;/p&gt;</description></item><item><title>A Prompt Is Not Yet a Program</title><link>https://aibussin.com/books/dspy-from-first-principles/02-chapter/</link><pubDate>Fri, 28 Aug 2026 10:05:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/02-chapter/</guid><description>&lt;p&gt;Chapter 1 ended with a signature and an uncomfortable amount of unfinished business. We had established that one prompt string was carrying the task semantics, the behavioral framing, the output schema, and the provider assumptions simultaneously, and that this is why three unrelated failures arrived looking identical.&lt;/p&gt;&#10;&lt;p&gt;The obvious response is to write a better string. That is the wrong move, and it is worth being precise about why.&lt;/p&gt;&#10;&lt;p&gt;A prompt is text:&lt;/p&gt;</description></item><item><title>Define the Contract</title><link>https://aibussin.com/books/dspy-from-first-principles/03-chapter/</link><pubDate>Fri, 28 Aug 2026 10:10:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/03-chapter/</guid><description>&lt;p&gt;Chapter 2 built a boundary and then exposed a weakness in it on purpose. We designed the most natural output field available — &lt;code&gt;rewritten_text&lt;/code&gt;, a replacement sentence — but failed to represent &lt;em&gt;&amp;ldquo;no edit needed&amp;rdquo;&lt;/em&gt; as an explicit decision.&lt;/p&gt;&#10;&lt;p&gt;The model can technically return the original sentence unchanged. What the contract cannot tell us is whether that identity output means deliberate restraint or an attempted rewrite that happened to reproduce the input. That distinction matters on six of the forty-four later fixture cases — about &lt;code&gt;13.6%&lt;/code&gt; of the corpus. Having a signature did not prevent the mistake.&lt;/p&gt;</description></item><item><title>Separate What From How</title><link>https://aibussin.com/books/dspy-from-first-principles/04-chapter/</link><pubDate>Fri, 28 Aug 2026 10:15:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/04-chapter/</guid><description>&lt;p&gt;Chapters 2 and 3 built a contract and said nothing about how a model should satisfy it. That was deliberate, and this chapter collects the payment.&lt;/p&gt;&#10;&lt;p&gt;Three objects must remain distinct throughout the comparison:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;task contract what valid behavior must preserve&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;execution strategy how the model attempts the task&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;evaluation how observed behavior is judged&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The signature holds the first fixed. &lt;code&gt;Predict&lt;/code&gt; and &lt;code&gt;ChainOfThought&lt;/code&gt; vary the second. The frozen fixture and metrics supply the third. If two of those change together, the result no longer tells us what the reasoning policy caused.&lt;/p&gt;</description></item><item><title>Build Programs From Programs</title><link>https://aibussin.com/books/dspy-from-first-principles/05-chapter/</link><pubDate>Fri, 28 Aug 2026 10:20:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/05-chapter/</guid><description>&lt;p&gt;Chapter 4 held the contract fixed and changed the execution policy. This chapter changes the architecture.&lt;/p&gt;&#10;&lt;p&gt;Real applications rarely consist of one call. An editorial system might want to work out what is wrong with a sentence, produce a candidate, assess whether that candidate is risky, and hand the whole package to something that stores it.&lt;/p&gt;&#10;&lt;pre class="mermaid"&gt;&#10; flowchart LR&#10; I[input] --&amp;gt; A[analyze]&#10; A --&amp;gt; W[rewrite]&#10; W --&amp;gt; S[assess]&#10; S --&amp;gt; R[return]&#10; &lt;/pre&gt;&#10; &lt;p&gt;That is the pipeline as pitched: four stages in a line, each feeding the next. Decomposition is one of the few architectural moves that is widely treated as obviously correct. Smaller pieces, clearer boundaries, better debugging — the diagram sells itself.&lt;/p&gt;</description></item><item><title>Your Model Is a Dependency</title><link>https://aibussin.com/books/dspy-from-first-principles/06-chapter/</link><pubDate>Fri, 28 Aug 2026 10:25:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/06-chapter/</guid><description>&lt;p&gt;Chapter 5 built a composed program and measured which of its stages earned their cost. Every number in that chapter, and in chapter 4 before it, has an unstated qualifier:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Which model executed the program?&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is not operational trivia. A language-model program is not fully described by its signatures and modules. Its behavior depends on what sits behind the LM boundary and how that thing is configured.&lt;/p&gt;&#10;&lt;p&gt;The model is therefore both a dependency and an experimental variable. Holding it fixed is often as important as swapping it:&lt;/p&gt;</description></item><item><title>Examples Are Experimental Data</title><link>https://aibussin.com/books/dspy-from-first-principles/07-chapter/</link><pubDate>Fri, 28 Aug 2026 10:30:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/07-chapter/</guid><description>&lt;p&gt;Chapter 6 gave the program a dependency boundary and, more importantly, evidence that identical recorded program state can still produce different outputs. We can now say what the program is, what ran it, and why a small score delta cannot be interpreted without repeated measurement.&lt;/p&gt;&#10;&lt;p&gt;What we still cannot do is say whether the program is any good, because we have nothing to measure it against.&lt;/p&gt;&#10;&lt;p&gt;Examples in DSPy are not decoration around the program. They are experimental data with several possible roles. Optimization consumes that evidence: training data may shape candidate state, development data may select among candidates, and holdout data must remain outside both processes if it is to support an independent claim.&lt;/p&gt;</description></item><item><title>You Cannot Optimize What You Cannot Measure</title><link>https://aibussin.com/books/dspy-from-first-principles/08-chapter/</link><pubDate>Fri, 28 Aug 2026 10:35:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/08-chapter/</guid><description>&lt;p&gt;Chapter 7 produced material. This chapter produces evidence.&lt;/p&gt;&#10;&lt;p&gt;That makes optimization possible. It does not yet make optimization safe.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;no metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → no systematic optimization&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → behavior can be optimized&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;wrong metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → the wrong behavior can be optimized systematically&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This chapter establishes the surface on which search can operate. Chapter 9 asks what that surface rewards when the metric becomes the target.&lt;/p&gt;&#10;&lt;pre class="mermaid"&gt;&#10; flowchart TD&#10; P[frozen program] --&amp;gt; E[evaluate]&#10; C[frozen cases] --&amp;gt; E&#10; M[metric] --&amp;gt; E&#10; E --&amp;gt; PC[per-case results]&#10; PC --&amp;gt; AG[aggregate result]&#10; PC --&amp;gt; FI[failure inspection]&#10; &lt;/pre&gt;&#10; &lt;p&gt;No optimization happens here. Note that the per-case results feed &lt;em&gt;two&lt;/em&gt; readings — the aggregate and the failure inspection — and this chapter&amp;rsquo;s finding comes from the second. The question is narrow:&lt;/p&gt;</description></item><item><title>When the Metric Becomes the Target</title><link>https://aibussin.com/books/dspy-from-first-principles/09-chapter/</link><pubDate>Fri, 28 Aug 2026 10:40:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/09-chapter/</guid><description>&lt;p&gt;Chapter 8 built a metric and found a defect in it by reading a table. Six cases where the correct answer is capped at 0.65, discovered by grouping per-case scores and noticing that a whole category sat on an exact number.&lt;/p&gt;&#10;&lt;p&gt;That was luck dressed as diligence. This chapter looks on purpose.&lt;/p&gt;&#10;&lt;p&gt;The urgency comes from what happens next. A DSPy optimizer does not want your program to be good. It wants your number to go up, and it will find every route to that outcome, including the ones you did not intend to leave open.&lt;/p&gt;</description></item><item><title>Compile the Program</title><link>https://aibussin.com/books/dspy-from-first-principles/10-chapter/</link><pubDate>Fri, 28 Aug 2026 10:45:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/10-chapter/</guid><description>&lt;p&gt;Chapter 9 gave the optimizer an explicit objective and showed a concrete limitation: v1 can assign very high scores to candidates that reverse important parts of the source meaning.&lt;/p&gt;&#10;&lt;p&gt;We are going to compile against that frozen metric anyway. Not because the limitation is acceptable, but because changing the objective now would change the experiment. v2 remains outside the main search loop as an independent semantic check.&lt;/p&gt;&#10;&lt;p&gt;Compilation in DSPy takes a program, training evidence, a metric, and an optimizer, and returns a candidate program state.&lt;/p&gt;</description></item><item><title>Few-Shot Optimization</title><link>https://aibussin.com/books/dspy-from-first-principles/11-chapter/</link><pubDate>Fri, 28 Aug 2026 10:50:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/11-chapter/</guid><description>&lt;p&gt;Chapter 10 produced a candidate and refused to say whether it was any good. This chapter says.&lt;/p&gt;&#10;&lt;p&gt;The mechanism is few-shot optimization, and the question is practical:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Can we improve the program by selecting better demonstrations, rather than rewriting the program ourselves?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;The answer, on our fixture, is the most interesting result in this book. In one run, the candidate appears to improve the metric we optimized against while damaging the metric we did not. Under seven paired fresh-process sessions, the apparent v1 gain collapses to essentially zero while the v2 loss remains large and systematic.&lt;/p&gt;</description></item><item><title>Optimize the Instructions</title><link>https://aibussin.com/books/dspy-from-first-principles/12-chapter/</link><pubDate>Fri, 28 Aug 2026 10:55:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/12-chapter/</guid><description>&lt;p&gt;Chapter 11&amp;rsquo;s optimizer could only select demonstrations. It never touched an instruction, and it never looked at the development set — which is why upgrading its metric changed nothing.&lt;/p&gt;&#10;&lt;p&gt;MIPROv2 does both.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;instruction candidates&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;demonstration candidates&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;development data&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;candidate program&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is a strictly larger search over a strictly larger space, evaluated against evidence the previous optimizer could not see. If chapter 11&amp;rsquo;s result was an artifact of a weak mechanism, this is where it gets corrected.&lt;/p&gt;</description></item><item><title>Let the Program Reflect</title><link>https://aibussin.com/books/dspy-from-first-principles/13-chapter/</link><pubDate>Fri, 28 Aug 2026 11:00:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/13-chapter/</guid><description>&lt;p&gt;Chapter 12 ended on a specific limitation. MIPROv2 selected a candidate that deleted &lt;code&gt;regularly&lt;/code&gt; from a regulated claim, and nothing in the pipeline could have told it otherwise, because the only signal it received was &lt;code&gt;0.967&lt;/code&gt;.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;candidate&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;0.967&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A number can rank candidates. It cannot say which word was the problem.&lt;/p&gt;&#10;&lt;p&gt;The next mechanism replaces the scalar with language:&lt;/p&gt;&#10;&lt;pre class="mermaid"&gt;&#10; flowchart LR&#10; P[prediction] --&amp;gt; M[metric]&#10; M --&amp;gt; SF[score + diagnostic feedback]&#10; SF --&amp;gt; RF[reflection]&#10; RF --&amp;gt; IM[instruction mutation]&#10; IM --&amp;gt; CA[candidate]&#10; CA --&amp;gt;|re-run| P&#10; &lt;/pre&gt;&#10; &lt;p&gt;GEPA is DSPy&amp;rsquo;s evolutionary, reflection-driven optimizer. It reads per-predictor feedback, uses a separate reflection model to propose instruction edits, maintains a population of candidate programs, and returns the best one it found. The engineering idea underneath is smaller than the API:&lt;/p&gt;</description></item><item><title>Change One Thing</title><link>https://aibussin.com/books/dspy-from-first-principles/14-chapter/</link><pubDate>Fri, 28 Aug 2026 11:05:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/14-chapter/</guid><description>&lt;p&gt;The previous three chapters ran three optimizers against a metric we had already proved was blind, and got what chapter 9 predicted: a search that pursued the objective faithfully and damaged the thing the objective could not see.&lt;/p&gt;&#10;&lt;p&gt;That result has an obvious interpretation and a correct one, and they are not the same.&lt;/p&gt;&#10;&lt;p&gt;The obvious interpretation is that the optimizers underperformed. The correct one requires an experiment, because &amp;ldquo;the objective was the problem&amp;rdquo; is a hypothesis and we had not tested it. Everything we knew was consistent with a second explanation — that DSPy&amp;rsquo;s optimizers do not do much on a small local model regardless of what you point them at.&lt;/p&gt;</description></item><item><title>Agents Are Programs Too</title><link>https://aibussin.com/books/dspy-from-first-principles/15-chapter/</link><pubDate>Fri, 28 Aug 2026 11:05:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/15-chapter/</guid><description>&lt;p&gt;Chapter 14 held the corpus, program, model, optimizer configurations and paired harness fixed, changed only the objective, and closed the editorial task. The prediction it left behind was that a reflective optimizer improves in proportion to how much its feedback can say — and that a task whose failures carry a test name, an expected value and a traceback should therefore be the book&amp;rsquo;s best case.&lt;/p&gt;&#10;&lt;p&gt;To measure that we need a task with those properties. Repository repair has them. But it introduces a limitation the editorial task never had.&lt;/p&gt;</description></item><item><title>Search, Memory and Long Context</title><link>https://aibussin.com/books/dspy-from-first-principles/16-chapter/</link><pubDate>Fri, 28 Aug 2026 11:10:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/16-chapter/</guid><description>&lt;p&gt;Chapter 15 measured an agent that had every advantage. Two files, one defect, a bounded tool surface, a budget it used well. It found both implicated files, held both halves of the bug in context, and still stopped one step short. That was a reasoning failure, and Chapter 17 exists to attack it.&lt;/p&gt;&#10;&lt;p&gt;This chapter is about the problem that sits &lt;em&gt;before&lt;/em&gt; the reasoning problem. The fixture had two files. A real repository has thousands, and a bounded tool surface does not tell the program which of them to read. Something has to decide what evidence the program even looks at, and that decision is not one mechanism. It is at least five, and they are routinely collapsed into a single word.&lt;/p&gt;</description></item><item><title>Search the Reasoning Space</title><link>https://aibussin.com/books/dspy-from-first-principles/17-chapter/</link><pubDate>Fri, 28 Aug 2026 11:25:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/17-chapter/</guid><description>&lt;h2 id="one-path-is-a-decision"&gt;One Path Is a Decision&lt;/h2&gt;&#10;&lt;p&gt;In Chapter 15, we built an agent. It acquired repository evidence safely. It inspected the files it was permitted to inspect. It produced a diagnosis.&lt;/p&gt;&#10;&lt;p&gt;It was wrong.&lt;/p&gt;&#10;&lt;p&gt;Not completely wrong—the agent scored &lt;code&gt;0.60&lt;/code&gt;. It found the relevant code. It identified that &lt;code&gt;service.py&lt;/code&gt; references &lt;code&gt;item.cost&lt;/code&gt;. It located the model definition. And then it concluded that &lt;code&gt;Item&lt;/code&gt; has no &lt;code&gt;cost&lt;/code&gt; attribute, missing the final connection: &lt;code&gt;service.py&lt;/code&gt; should use &lt;code&gt;item.price&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Don't Let the Optimizer Cheat</title><link>https://aibussin.com/books/dspy-from-first-principles/18-chapter/</link><pubDate>Fri, 28 Aug 2026 11:15:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/18-chapter/</guid><description>&lt;h2 id="opening--the-optimizer-doesnt-have-to-cheat-intentionally"&gt;Opening — The Optimizer Doesn&amp;rsquo;t Have to Cheat Intentionally&lt;/h2&gt;&#10;&lt;p&gt;In Chapters 15 and 16 we expanded the program&amp;rsquo;s access. The agent could search repositories, retrieve memories, and adaptively select evidence. Chapter 17 then added search over reasoning states, showing that tree search can allocate inference-time compute toward promising paths.&lt;/p&gt;&#10;&lt;p&gt;Now the question becomes urgent: once a program can search, remember, and adaptively select both evidence and reasoning, how do we prove that it did not obtain information it was never supposed to see?&lt;/p&gt;</description></item><item><title>From Experiment to Production</title><link>https://aibussin.com/books/dspy-from-first-principles/19-chapter/</link><pubDate>Fri, 28 Aug 2026 11:20:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/19-chapter/</guid><description>&lt;p&gt;Chapter 18 ended with an information boundary we could finally audit: the firewall blocked all 288 constructed attacks in its deterministic train/dev suite. That does &lt;strong&gt;not&lt;/strong&gt; make the downstream metric correct; it means the tested result was not produced by the forbidden evidence paths we attacked. Chapter 18 also said the obvious next thing — even an admissible offline result is not a production version.&lt;/p&gt;&#10;&lt;p&gt;This chapter closes two gaps between those two facts.&lt;/p&gt;</description></item><item><title>Build a Self-Improving Engineering Program</title><link>https://aibussin.com/books/dspy-from-first-principles/20-chapter/</link><pubDate>Fri, 28 Aug 2026 11:25:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/20-chapter/</guid><description>&lt;p&gt;We can now assemble the book.&lt;/p&gt;&#10;&lt;p&gt;Every prior chapter added one mechanism because the system before it had exposed a specific limitation: a contract (Chapter 3), a measured baseline (Chapter 8), an attacked metric (Chapter 9), three optimizers and an objective swap (Chapters 11–14), tools and bounded environment access (Chapter 15), evidence selection (Chapter 16), reasoning-path search (Chapter 17), an experimental firewall (Chapter 18), and a promotion boundary with rollback (Chapter 19). This chapter does not introduce another mechanism. It asks whether the ones we have compose into a single loop without any stage quietly promoting its own output.&lt;/p&gt;</description></item><item><title>Beyond the Book: Production DSPy Systems (Appendix)</title><link>https://aibussin.com/books/dspy-from-first-principles/21-chapter/</link><pubDate>Thu, 03 Sep 2026 11:00:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/21-chapter/</guid><description>&lt;p&gt;Chapter 20 assembled the book&amp;rsquo;s mechanisms into one loop. This appendix does something smaller and more specific: it shows three of those mechanisms already running in a production codebase that was not written for this book, and it holds itself to a stricter evidence rule than the chapters do.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Evidence rule for this appendix.&lt;/strong&gt; There are two tiers here. The &lt;strong&gt;mechanics&lt;/strong&gt; around each language-model call — UCT selection, the budget gate, the trace cache, the champion-update rule, the output parser — are deterministic software, and a committed fixture (&lt;code&gt;experiments/dspy-from-first-principles/ch21_beyond_book/&lt;/code&gt;) exercises them with no LM and no DSPy import. Those results are measured. The &lt;strong&gt;kernels themselves&lt;/strong&gt; — the production modules the excerpts are drawn from — are &lt;em&gt;selection evidence from one codebase&lt;/em&gt;, not a generalisation. Their full source lives in the &lt;a href="https://github.com/ernanhughes/stephanie"&gt;Stephanie repository&lt;/a&gt;, linked below and pinned to &lt;code&gt;main&lt;/code&gt; at revision &lt;code&gt;32d02ee&lt;/code&gt; (2026-01-06).&lt;/p&gt;</description></item></channel></rss>