← DSPy From First Principles

Don't Let the Optimizer Cheat

Build an experimental firewall for DSPy optimization so tools, memory, feedback, retrieval, and holdout evidence cannot quietly leak answers.

Opening — The Optimizer Doesn’t Have to Cheat Intentionally

In Chapters 15 and 16 we expanded the program’s access. The agent could search repositories, retrieve memories, and adaptively select evidence. Chapter 17 then added search over reasoning states, showing that tree search can allocate inference-time compute toward promising paths.

Now the question becomes urgent: once a program can search, remember, and adaptively select both evidence and reasoning, how do we prove that it did not obtain information it was never supposed to see?

That power threatens the experiment.

The central question:

How can an optimization experiment appear valid while the program or optimizer has quietly gained access to the answer?

Cheating is often accidental. A field named validation_result slips into inputs. A memory tool returns the historical accepted patch. A GEPA feedback string includes the exact gold answer. A developer inspects holdout failures, changes instructions, and evaluates on the same holdout again.

None of these require malicious intent. They require only an information boundary that is not enforced.

The experiment needs a firewall.


18.1 Leakage Is an Information-Flow Failure

Before listing checks, we need a threat model. What are we protecting, and from whom?

Leakage is not merely a bad field name. It is a violation of information flow: protected information crosses a boundary that should have been closed.

Let’s define the actors and the boundaries.

    graph TD
    subgraph Protected Information
        P1[Reference / Outcome labels]
        P2[Future states / holdout]
        P3[Historical accepted solutions]
        P4[Answer-bearing tool capabilities]
    end

    subgraph Consumers
        C1[Task program at inference]
        C2[Optimizer / DSPy compiler]
        C3[Retriever / memory]
        C4[Tool runtime]
        C5[Developer during tuning]
    end

    P1 -->|must not flow during training/inference| C1
    P1 -->|must not flow into candidate generation| C2
    P2 -->|must not be visible before evaluation| C2
    P3 -->|must not be retrieved by memory| C3
    P4 -->|must not be available if not permitted| C4
    P2 -->|must not guide manual changes| C5

    C1 -->|allowed: decision-time inputs| OK1[Admissible evidence]
    C2 -->|allowed: train/dev evidence| OK2[Optimizer-visible set]
    C3 -->|allowed: admissible corpus| OK3[Sanitized memory]
    C4 -->|allowed: pinned tools/revisions| OK4[Tool manifest]
    C5 -->|allowed: only pre-freeze diagnostics| OK5[Development window]
  

The principle:

Leakage is an information-flow violation. Not merely a bad field name.

We must design boundaries, and then enforce them with mechanisms that cannot be bypassed by accident.

Actor / subsystemWhat it may seeWhat it must not see
task programdecision-time inputsreference/outcome
optimizertrain/dev evidence permitted by protocolholdout
metriclabels/reference as required for evaluationmust not feed them back to candidate
retrieveradmissible corpusfuture/holdout/outcome sources
tool runtimepinned environmentlater repository state
memoryadmissible historical evidenceoutcome-derived content for current case
human developerdevelopment diagnosticsuntouched final holdout before freeze

We’ll build a firewall stack that enforces these boundaries.


18.2 A Split Is Not a Firewall

The simplest protection is a train/dev/holdout split. We keep some cases out of the optimizer’s visible set, and evaluate the final candidate on them.

That is necessary but insufficient.

A split controls which cases are visible during optimization. It does not control what information about those cases flows into the candidate.

A split can be defeated by:

  • A field that accidentally contains the label.
  • A retrieval system that returns a holdout case because it shares source lineage.
  • A tool that can read future commits where the solution exists.
  • A feedback signal that includes the reference answer.

A split is a policy. It assumes that if you don’t name a case as “holdout,” the optimizer can’t see it. But information leaks through many paths that are not named “holdout”.

We need explicit firewalls that inspect the actual information flow, not just the IDs.

The rest of this chapter builds a six-layer firewall that does exactly that.


18.3 The Six-Layer Experimental Firewall

The firewall is a stack of independent checks. Each layer catches something that earlier layers cannot.

    graph TD
    A[Candidate / Optimizer Request] --> B[1. Schema Firewall]
    B --> C[2. Content Firewall]
    C --> D[3. Lineage Firewall]
    D --> E[4. Temporal Firewall]
    E --> F[5. Capability Firewall]
    F --> G[6. Split Firewall]
    G --> H{Admissible?}
    H -->|Yes| I[Run Execution]
    H -->|No| J[Reject / Quarantine]
  

Each layer answers a different question:

  1. Schema: Are there fields or keys that should never appear in the input payload?
  2. Content: Does any input value match (or closely resemble) protected content, regardless of key?
  3. Lineage: Does the evidence come from the same source as a holdout case, even under a different ID?
  4. Temporal: Is the evidence from a time after the case’s decision point (e.g., future commit)?
  5. Capability: Does the runtime expose tools or permissions that could access protected information, even if they are not called?
  6. Split: Do any explicitly designated holdout IDs leak into the optimizer-visible set?

We’ll explore each in turn, with concrete implementations.


18.4 Layer 1 — Schema and Outcome Fields

The most obvious leak: a field that should be output-only appears in inputs.

In the CoCoder project, the corpus layer defines outcome keys such as:

OUTCOME_KEYS = {
    "reference_rewrite",
    "human_accepted",
    "validation_result",
    "metric_score",
    "promotion_decision",
}

A schema firewall checks that none of these keys appear in the input payload:

def assert_no_leakage(inputs: dict) -> None:
    leaked = sorted(key for key in inputs if key in OUTCOME_KEYS)
    if leaked:
        raise ValueError(f"input contains outcome leakage: {', '.join(leaked)}")

But this only catches the key name. It does not catch the same gold text under an innocent key like retrieval_hint.

Consider:

bad_example = dspy.Example(
    sentence=row["sentence"],
    editorial_goal=row["editorial_goal"],
    local_context=row["local_context"],
    reference_rewrite=row["reference_rewrite"],
).with_inputs(
    "sentence",
    "editorial_goal",
    "local_context",
    "reference_rewrite",
)

The model now receives the reference answer. Schema firewall blocks it because reference_rewrite is in OUTCOME_KEYS.

But if someone renames it:

bad_example = dspy.Example(
    sentence=row["sentence"],
    editorial_goal=row["editorial_goal"],
    local_context=row["local_context"],
    retrieval_hint=row["reference_rewrite"],  # disguised
).with_inputs(
    "sentence", "editorial_goal", "local_context", "retrieval_hint"
)

Schema check passes. The answer is still in the inputs.

So we need Layer 2.


18.5 Layer 2 — Content and Paraphrase

The content firewall inspects values, not keys. It checks whether any input field contains exact protected content, or content that is highly similar.

For exact match:

PROTECTED_VALUES = {
    "exact_reference_string_1",
    "exact_reference_string_2",
    # ...
}

def assert_no_exact_content_leak(inputs: dict) -> None:
    for key, value in inputs.items():
        if isinstance(value, str) and value in PROTECTED_VALUES:
            raise ValueError(f"content leak in field '{key}'")

For n-gram containment (to catch partial leakage):

def ngrams(text: str, n: int = 5) -> set[str]:
    tokens = text.split()
    return {" ".join(tokens[i:i+n]) for i in range(len(tokens) - n + 1)}

def assert_no_ngram_leak(inputs: dict, protected_values: set[str], threshold: float = 0.8) -> None:
    for key, value in inputs.items():
        if not isinstance(value, str):
            continue
        input_ngrams = ngrams(value)
        for pv in protected_values:
            pv_ngrams = ngrams(pv)
            overlap = len(input_ngrams & pv_ngrams) / max(1, len(pv_ngrams))
            if overlap > threshold:
                raise ValueError(f"ngram overlap between '{key}' and protected value")

For fuzzy semantic similarity, we can use embeddings. But embeddings are heuristic, not authoritative:

from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("all-MiniLM-L6-v2")

def semantic_similarity(a: str, b: str) -> float:
    emb_a = model.encode(a)
    emb_b = model.encode(b)
    return float(util.cos_sim(emb_a, emb_b))

# Use only as a flag for manual review
def flag_semantic_similarity(inputs: dict, protected_values: set[str], threshold: float = 0.82):
    flags = []
    for key, value in inputs.items():
        if not isinstance(value, str):
            continue
        for pv in protected_values:
            sim = semantic_similarity(value, pv)
            if sim > threshold:
                flags.append((key, sim))
    return flags  # raise or quarantine

Important: A semantic similarity threshold is not a security guarantee. It can produce false positives (legitimate texts that are related to the topic) and false negatives (paraphrased hints that are not close in embedding space).

We treat semantic detection as a flag for audit, not an automatic reject. Exact content and high n-gram overlap are stronger evidence.

What the measured suite actually enforced. The 288-attack run later in this chapter used the exact-match and substring checks only. The n-gram-overlap ratio and the embedding flag above are the design this layer should grow into; they were not part of the enforced firewall when the numbers were collected, which is why every residual risk in section 18.12 is a paraphrase or abstraction that exact-plus-substring cannot see.


18.6 Layer 3 — Lineage and Duplicate Sources

A duplicate source can defeat both schema and content checks. If you have a holdout case with ID ed-004 derived from document X, and the optimizer’s training set contains document X under a different ID, the content may be paraphrased but the source lineage is the same.

Lineage firewall checks provenance: where did this example come from? Does it share a source with a holdout case?

@dataclass
class CaseLineage:
    case_id: str
    source_document_id: str | None
    source_url: str | None
    source_hash: str | None
    split: str  # "train", "dev", "holdout"

We group by source:

def assert_no_lineage_overlap(
    candidate_manifest: dict, holdout_lineages: set[tuple[str, str, str]]
) -> None:
    # candidate_manifest contains list of case lineages visible to optimizer
    for lineage in candidate_manifest["lineages"]:
        key = (lineage["source_document_id"], lineage["source_url"], lineage["source_hash"])
        if key in holdout_lineages:
            raise ValueError(f"lineage overlap with holdout: {lineage['case_id']}")

If two cases come from the same source document, even with different IDs and paraphrased content, they are too close. The optimizer may learn from the holdout’s sibling and generalize in a way that is not legitimate.


18.7 Layer 4 — Time-Travel Leakage

Repository agents can inspect git history. If a case is pinned at a certain commit, but the agent can run git log, git show, or git diff to see later commits, it might discover the solution that was committed after the case was created.

issue occurs at commit A

A ---- B ---- C ---- D
^                 ^
case revision     fix committed here

The temporal firewall pins the repository state and restricts access to history beyond the case’s decision point.

class RepositoryPinner:
    def __init__(self, allowed_revision: str, allowed_history_end: str | None = None):
        self.allowed_revision = allowed_revision
        self.allowed_history_end = allowed_history_end or allowed_revision

    def assert_allowed_command(self, command: str) -> None:
        # Block any command that might reveal future history
        dangerous_commands = [
            "git log",
            "git show",
            "git diff",
            "git checkout",
            "git branch",
        ]
        if any(dc in command for dc in dangerous_commands):
            raise ValueError(f"command '{command}' may reveal future or unpinned history")
        # Also check that no command references a revision newer than allowed_history_end
        # ...

But the better approach is to run the agent in a pinned worktree where the repository is physically checked out at the allowed revision and no future commits exist. Then even if the agent tries git log, it only sees history up to that point.

The key insight: enforce at the tool layer, not by instruction. If the agent is told “don’t look at future commits” but the tool can see them, that’s a vulnerability.


18.8 Layer 5 — Capability Surfaces

Even if the agent never calls a dangerous tool, the fact that the tool is available changes the experimental protocol. If the agent could have accessed an answer-bearing file but chose not to, we cannot prove it didn’t.

So the firewall should audit the capability surface before execution.

@dataclass
class CapabilityManifest:
    allowed_tools: set[str]
    allowed_repo_revision: str
    network_allowed: bool
    shell_allowed: bool
    writable_paths: set[str]

Before running, we compare the declared expected surface with the actual runtime configuration:

def assert_capability_compliance(
    declared: CapabilityManifest,
    actual: CapabilityManifest,
) -> None:
    # Actual must be a subset of declared
    if not actual.allowed_tools.issubset(declared.allowed_tools):
        extra = actual.allowed_tools - declared.allowed_tools
        raise ValueError(f"unexpected tools available: {extra}")
    if actual.allowed_repo_revision != declared.allowed_repo_revision:
        raise ValueError("repository revision mismatch")
    if actual.network_allowed and not declared.network_allowed:
        raise ValueError("network unexpectedly enabled")
    if actual.shell_allowed and not declared.shell_allowed:
        raise ValueError("shell unexpectedly enabled")
    if not actual.writable_paths.issubset(declared.writable_paths):
        extra = actual.writable_paths - declared.writable_paths
        raise ValueError(f"unexpected writable paths: {extra}")

If a run is supposed to be read-only, the runtime should not even mount writable paths. If network is disallowed, the container should have no network interface. This is a structural guarantee, not a runtime check.


18.9 Layer 6 — Split and Holdout Isolation

Finally, we ensure that explicit holdout IDs never enter the optimizer’s visible set, either as demonstrations or as visible cases.

def assert_holdout_not_visible(candidate_manifest: dict, holdout_ids: set[str]) -> None:
    demos = set(candidate_manifest.get("demonstration_case_ids", []))
    visible = set(candidate_manifest.get("optimizer_visible_case_ids", []))
    leaked = sorted((demos | visible) & holdout_ids)
    if leaked:
        raise ValueError(f"holdout cases were visible: {', '.join(leaked)}")

This is the most basic check, and it should be applied to every artifact that touches the optimizer: prompts, few-shot examples, retrieval indexes, tool outputs, and logs.

But holdout exhaustion is a more subtle violation. If we inspect holdout failures and then modify the system, the holdout becomes a development set. We need to track that:

class HoldoutTracker:
    def __init__(self, holdout_ids: set[str]):
        self.holdout_ids = holdout_ids
        self.exposed_to_developer: set[str] = set()

    def mark_developer_visible(self, case_id: str) -> None:
        if case_id in self.holdout_ids:
            self.exposed_to_developer.add(case_id)

    def assert_no_exhaustion(self) -> None:
        if self.exposed_to_developer:
            raise ValueError(
                f"holdout cases have been exposed to developer: {self.exposed_to_developer}. "
                "They are no longer a valid holdout."
            )

Once a holdout case’s outcome influences a design decision, it is no longer independent. It must be moved to the development set and replaced with a fresh case.


18.10 When the Firewall Blocks the Correct Answer

The firewall layers we’ve built are designed to block leakage. But they can also block legitimate behavior.

In our editorial rewrite task, sometimes the correct answer is no change. The sentence is already fine. The reference rewrite is identical to the original.

input sentence: "The service should use item.price rather than item.cost."
reference_rewrite: "The service should use item.price rather than item.cost."

The firewall rule we had was:

the program must not receive/emit the protected reference value

But if the reference value equals the original sentence, then a clean correct answer is indistinguishable from a leaked answer. The firewall would reject a legitimate case.

This is a restraint failure: the firewall cannot distinguish between the model copying the answer and the model producing the only correct output.

This connects directly to earlier chapters:

  • In Chapter 2, the contract didn’t explicitly represent “no edit needed”.
  • In Chapter 8, the metric scored an unchanged correct answer as 0.65.
  • In Chapter 9, the representation problem appeared at the metric layer.

Now it appears at the security layer: a security invariant cannot be defined independently of the task semantics.

The resolution: for cases where the correct answer is identical to a protected value, we must either:

  1. Remove them from the evaluation set, or
  2. Mark them as identity-reference cases and allow the firewall to pass if the output is identical to the input but not copied from any external source.

This requires tracking whether the program had access to the reference value at all. If it didn’t, and it still produced the same string, that’s not leakage—it’s correct.

So the firewall must be context-aware: it must know whether the protected content could have been seen. If not, matching the reference is not evidence of leakage.

This is a powerful reminder: security is not a boolean property; it is a property relative to the information available.


18.11 Attack the Firewall

We didn’t just build the firewall and trust it. We attacked it.

The deterministic attack suite now covers 32 eligible cases across the editorial task. For each case, we constructed 9 attack variants that attempt to smuggle forbidden evidence into the program’s inputs or the optimizer’s visible set.

Total executions: 32 × 9 = 288.

The attack matrix:

AttackWhat crossed boundaryExpected firewall layer
outcome_keyexplicit outcome key (reference_rewrite)schema
innocent_key_gold_valuegold answer under innocent key (retrieval_hint)content
gold_feedbackexact gold feedback string in GEPAcontent
duplicate_sourcesame source document as holdout, new IDlineage
future_revisionevidence from a later commitrevision
accepted_memoryhistorical accepted patch retrievedcontent/memory
answer_bearing_tooltool capable of reading answer filecapability
holdout_manifestholdout ID in optimizer-visible listsplit
identity_referencecorrect no-change case where output equals referencerestraint (special handling)

All 288 attacks were blocked by the firewall. No clean payload produced a false positive.

But the crucial number: 224 of the attack executions would have inflated the score if the evidence had been admitted. The mean reference-contamination effect was +0.3856.

without firewall
      ↓
224 / 288 attacks produce apparent improvement

with firewall
      ↓
288 / 288 inadmissible paths rejected

This demonstrates both that the firewall is working and that the temptation to cheat (accidentally or otherwise) is real.

The identity_reference attacks are particularly interesting. They don’t provide the answer; they test whether the firewall incorrectly blocks a legitimate no-change case. In our suite, we had to special-case them: the firewall was adjusted to allow identity-reference cases when the program had no access to the protected value. This illustrates the restraint failure we discussed in 17.10.


18.12 False Negatives

Despite the success, no firewall is perfect. We must consider what could still slip through.

  • Paraphrase below threshold: A summary of the gold answer that avoids exact n-grams and is semantically close but below the embedding threshold may pass content checks.
  • Abstracted hints: “The bug is in service.py, look at how item fields are used” might be enough to point the agent to the solution without containing protected text.
  • Derived artifacts: An optimizer that sees evaluation scores can infer which directions are promising, even without seeing the answer. This is a weaker form of leakage but still contaminates the result.
  • Future state through side channels: Even without git history, the existence of a file with a certain modification time could hint at the solution.
  • Tool logs: The agent might not call a dangerous tool, but the logs of previous runs (if not cleared) could contain the answer.

Therefore, we should pair the firewall with audit logging and red teaming. The firewall is a first line of defense, not a guarantee.


18.13 Multi-Agent Leakage

Information-flow guarantees must be transitive. Consider a multi-agent setup:

Agent A
 sees forbidden evidence
      ↓
sends summary to Agent B
      ↓
Agent B
 never directly sees forbidden source
      ↓
uses leaked information

Agent B’s own input may be clean, but the information has already been contaminated by Agent A. A simple per-agent firewall would miss this.

We need taint propagation. Each message between agents carries lineage:

@dataclass
class MessageLineage:
    source_case_ids: set[str]
    repository_revisions: set[str]
    tool_calls: list[str]
    contains_outcome_evidence: bool = False

If Agent A consumed a forbidden source, its output becomes tainted. When that output is sent to Agent B, the taint propagates.

def propagate_taint(msg: dict, tainted: bool) -> dict:
    msg["tainted"] = tainted
    return msg

# Later, when evaluating Agent B's candidate:
if any(msg.get("tainted", False) for msg in agent_b_inputs):
    raise ValueError("candidate used tainted information")

This extends the firewall from a single process to a graph of information flow.


18.14 Firewall Invariants

We can now state the invariants concisely:

  1. Schema Invariant: No field in the program input may have a name in the forbidden outcome-key set.
  2. Content Invariant: No program input value may exactly match or contain high n-gram overlap with protected content.
  3. Lineage Invariant: No visible case may share the same source lineage as a holdout case.
  4. Temporal Invariant: No evidence from a repository revision later than the case’s pinned revision may be accessed.
  5. Capability Invariant: The runtime’s actual tool and permission surface must be a subset of the declared surface.
  6. Split Invariant: No holdout ID may appear in any optimizer-visible manifest.
  7. Restraint Invariant: For identity-reference cases, output equal to reference is not flagged if the program could not have seen the reference.
  8. Transitivity Invariant: If any upstream agent consumed inadmissible evidence, all downstream outputs are tainted.

These invariants form the experimental firewall. They are implemented as explicit checks in the experiment runner, not as policies in a document.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
Holdout score keeps improving after manual tuningHoldout is now devReview change history after first holdout runFreeze a new final test set
Innocent-looking field contains the gold answerFirewall checks keys but not valuesCompare payload values against protected outcomesAdd content/value leakage checks
Agent finds exact historical patchMemory/tool leaked outcomeInspect retrieved case IDs, source lineage, and fieldsFilter memory by split, lineage, and field policy
GEPA produces reference-specific instructionsFeedback leaked gold answerRead feedback records and protected-value matchesGive failure reason, not full answer
Candidate demos include holdout IDsSplit validation missingCompare demo IDs with holdout IDsReject artifact and rerun
Future repository file appears in evidenceTool used wrong revisionCheck repository revision in tool recordsPin worktree/revision per case
Agent never calls the dangerous tool, but it was availableCapability surface itself is broader than the protocolAudit declared tools before executionReject runs with forbidden capabilities
Correct no-change case gets blockedIdentity-reference restraint failureLook for cases where output == input == referenceSpecial-case identity-reference cases

Conclusion

We gained an experimental firewall and then attacked it.

The deterministic suite executed 288 attacks across 32 eligible cases. All 288 were blocked. Clean payloads produced zero false positives, and every expected detection contract was satisfied.

The attacks also demonstrated why the firewall matters. 224 contaminated executions would have improved the apparent score, by an average of about 0.3856, if the forbidden evidence had been admitted.

The strongest lesson is not merely “remove reference_rewrite from inputs.”

A split is not a firewall, and a forbidden-key list is not a firewall.

Gold information can arrive under an innocent key, inside feedback, from a duplicate source, through a future revision, via memory, or simply because the tool surface contains an answer-bearing capability. Preventing leakage therefore requires schema, content, lineage, revision, capability, and split checks working together, with additional attention to identity-reference cases and transitive taint.

We removed the assumption that a split alone prevents leakage.

Now we can produce a candidate we trust enough to consider deploying. But an offline win is not yet a production version.

In the next chapter, we bridge from experiment to production: what changes when the program must run on live data, with real users, and without the experimenter’s safety rails.