← Applied AI

If There's Any Doubt, It's Deterministic

The governing rule of the book: a component belongs on the deterministic side unless it demonstrably cannot be, and the doubt itself is the evidence. Plus the measured fact that temperature=0 does not make a model deterministic.

Part 1 — Where You Stand

Seven operations, one of them intelligent

Take the paragraph review from Chapter 1 and write down everything the process actually has to do. Not the prompt — the process.

  1. Select the paragraph to review.
  2. Decide whether it has changed since the last review.
  3. Decide whether there is budget left for a call.
  4. Identify claims in the paragraph that need supporting evidence.
  5. Extract the flagged sentence and its position.
  6. Check that a cited source resolves and that its year matches the bibliography.
  7. Decide whether to write the change to the file.

Six of those have exactly one correct answer, computable without a model. Selection is an index lookup and change detection is a hash comparison. Budget is arithmetic; extraction is parsing. Source checking resolves the request and compares it against the bibliography. The write decision evaluates policy.

One of them — step 4 — requires reading a sentence and forming a judgment about what would count as support for it. That is the operation you cannot specify well enough to implement.

Now look at how this is usually built. One prompt: “Review this paragraph, check the citations, and tell me if it’s ready to publish.” Seven operations, all seven routed through the stochastic component, and now the budget arithmetic can be wrong.

Which parts of a process are permitted to be unreliable, and who decides?

The rule

A component belongs on the deterministic side unless it demonstrably cannot be. If you are in doubt about which side it belongs on, that doubt is itself the evidence: it belongs on the deterministic side.

The second sentence sounds like a rhetorical flourish. It is better read as a strong default: before paying for stochastic generation, try to state the operation as explicit inputs, rules, and state transitions.

The important distinction is between specifying the operation and specifying a check on its output. If the operation itself can be computed from explicit state and rules, implement it deterministically. If you cannot write the generator but you can write an adequate check, that does not make the generator deterministic; it gives you the productive architecture this book uses repeatedly: stochastic proposal, deterministic verification.

So the operational test is:

Can this operation itself be computed from explicit state and rules? If not, can its result still be checked mechanically?

The first question decides whether a model is needed. The second decides how tightly a model can be bounded.

The asymmetry that makes any of this possible

Very often you cannot write the solution but you can write the check. You cannot write a function that repairs a broken parser, but you can run the test suite. You cannot write a function that produces a good paragraph review, but you can check that every sentence the review flagged actually exists in the paragraph, that every source it cites resolves, and that no source postdates the claim.

Where that check is substantially cheaper and more reliable than generation, applied AI gets leverage from the asymmetry. If checking an answer is itself open-ended, delayed, or adversarial, the verifier does not magically become deterministic just because the generator is stochastic.

Programs make the same division concrete. PAL has a language model write intermediate programs and hands them to an interpreter to execute, so the model proposes and exact machinery computes (Gao et al., 2023). Chapter 4 looks at benchmarks built deliberately on this asymmetry. The bound travels with it: the asymmetry holds where someone has manufactured an adequate cheap check. Where verification is delayed, subjective, or adversarial — fraud detection, in Chapter 4 — it does not.

So the split is not “model versus code.” It is:

    flowchart LR
    subgraph DET1["deterministic"]
        direction TB
        A["select · budget · policy"]
        B["context assembly<br/>explicit, hashed"]
    end
    subgraph STO["stochastic — confined"]
        M["model call<br/><i>proposes</i>"]
    end
    subgraph DET2["deterministic"]
        direction TB
        P["parse constrained output"]
        V["verify<br/><i>tests · source checks · reproduction</i>"]
        D["decide · act · record"]
    end
    A --> B --> M --> P --> V --> D
    D -.->|"unverified: retry, escalate, or stop"| A
  

One stochastic box. Deterministic machinery surrounds it. Proposals and failed checks may still be preserved as observations, but a proposal does not become an accepted effect merely because the model produced it.

Its unreliability is the product

Here is where the rule stops being a counsel of caution and becomes a design principle.

A deterministic function can compute answers from spaces far larger than anything you enumerated by hand. What makes it deterministic is narrower: given the same relevant state and inputs, the specified procedure produces the same result.

A model earns its place when the operation needs a proposal that you cannot economically produce with such a procedure. That proposal may contain something you did not anticipate, and the same variance that creates useful alternatives can also create failures.

This yields a practical heuristic:

If a different answer on re-run would be a defect, first look for a deterministic implementation.

The heuristic is not a proof. Sometimes generation is hard even when verification is exact; program repair behind a test suite is the obvious example. In that case the model may remain useful as a proposer while the deterministic check decides whether the proposal survives.

The same caution applies in the other direction. Multiple draws are useful only if their variation buys verified coverage. Part 5 tests that rather than assuming that different seeds, prompts, or models fail independently.

You cannot configure your way onto the deterministic side

The obvious objection: fine, but I’ll set temperature=0 and treat the call as deterministic.

That option is not available, and the reason is more interesting than the usual “floating point is fuzzy” hand-wave.

Horace He and colleagues at Thinking Machines Lab sampled 1,000 completions at temperature 0 from Qwen3-235B-A22B-Instruct-2507, same prompt, 1,000 tokens each. They got 80 unique completions. The most frequent one appeared 78 times. Every completion was identical for the first 102 tokens; divergence began at token 103, where 992 completions said “Queens, New York” and 8 said “New York City” (He et al., 2025).

Thinking Machines Lab makes a more specific diagnosis than the usual “GPU scheduling is random” explanation. In the inference stack they studied, the forward-pass kernels were run-to-run deterministic but not batch-invariant: changing batch size changed reduction order, which changed floating-point results. Replacing those kernels with batch-invariant versions made all 1,000 completions identical.

The architectural consequence is narrower and more useful. In a serving system with non-batch-invariant kernels, an individual request can depend on batch size, and batch size can depend on server load that the caller neither controls nor normally observes. Temperature zero therefore does not give the caller a reproducibility guarantee.

Yuan and colleagues found a related reproducibility problem under controlled local inference: changing evaluation batch size, GPU count, or GPU version altered generated responses under greedy decoding, with reasoning models amplifying small numerical differences into divergent trajectories.

Under bfloat16 with greedy decoding, DeepSeek-R1-Distill-Qwen-7B showed up to 9% variation in accuracy and 9,000 tokens of difference in response length attributable to GPU count, GPU type, and batch size (Yuan et al., 2025). Their mitigation, LayerCast, keeps weights in 16-bit but performs computation in FP32 — a change to the inference stack, not to an API parameter.

Bound these results properly. Both concern open-weight models and controlled inference stacks; neither measures the exact behavior of the hosted endpoints this book calls. The result the architecture needs is simply that greedy decoding removes sampling randomness, not every source of run-to-run variation visible to a caller.

The consequence for evaluation follows, and Song and colleagues make the broader point directly: evaluating a model on one output per example hides performance variability (Song et al., 2024). One run is a sample, not a measurement — which is why the experiments in Part 5 report repeated draws, and why Chapter 27 spends an entire chapter on a replication in which a promising signal did not survive.

So determinism is not a flag. You do not move a component to the deterministic column by configuring it. You move it by replacing it with something that computes the answer.

Classifying an operation

The decision procedure, applied to the seven steps we started with:

QuestionIf yesReview-pipeline example
Can explicit state and rules compute the operation?Deterministic. Implement it.Has the paragraph changed? (hash)
Is the allowed decision space enumerable in advance?Deterministic policy, not a prompt.Write, escalate, or stop?
Can you write an adequate check but not the generator?Stochastic proposal, deterministic verifier.Propose a repair; run the tests.
Do you need a proposal from a space you cannot economically enumerate?Stochastic, confined.Which claims might need evidence?
Would competent people reasonably disagree?Treat it as judgment, not proof that a model should decide.Is this phrasing adequately hedged?
Would variance on re-run be a defect?Look for a deterministic implementation first.Extracting a known sentence span.

The last row catches a common systems mistake, but the interface matters. Parsing constrained output against a known schema is deterministic. Trying to recover structure from arbitrary free-form prose may not be. The better fix is usually to constrain the upstream model output, then parse it mechanically, rather than adding a second model call to interpret the first.

The three deterministic rungs below reuse the paragraph review’s own steps. Assume the process has already stored the paragraph hash from the review, the spend so far, and the amount it is willing to reserve for one more call:

Chapter 28 later measures one small instance of the rule. In one frozen forty-item synthetic workload, ten inputs are already structured, with lines such as CDN_TTL_SECONDS=86400. The deterministic rule resolves all ten without a model call. In the top-first arm, the strong model resolves nine of those ten; on S10 it returns the bare digits 86400, the grounding check cannot verify the required duration expression, and the item goes to a person. The claim stays that size: on those ten mechanically specified items in that run, the rule added no model failure mode and the model did.

Where the rule breaks

A rule stated this firmly needs its own failure mode, and this one has a sharp one.

Determinism is cheap when the specification is cheap. When the specification is the dominant cost — open-ended natural language, a long tail of formats, judgments that resist enumeration — a hand-built rule cascade is not reliability. It is four hundred lines of brittle regex that someone maintains forever, fails silently on input twelve months from now, and is less trustworthy than a model call with a verifier behind it. “If in doubt, deterministic” does not license building a bad expert system in 2026.

When a deterministic step starts getting hard, check this first:

When a deterministic step becomes difficult, inspect the upstream interface before adding intelligence.

Extracting a flagged sentence from free-form model prose is genuinely miserable. That does not by itself show that extraction needs a model. If the generator can instead emit a schema with a quoted span or offsets, extraction becomes mechanical. When that redesign is available, adding a second model call to parse the first one’s prose expands the stochastic surface unnecessarily.

The honest boundary, then: the rule is a strong default with a clear override. Override it when a deterministic specification is genuinely impractical relative to the value it would provide — and record that choice, because a later model, interface, or workload may move the boundary.

What this buys you in money

Cost is not a side note in this book; it is a structural constraint, and it points the same way.

Deterministic operations are not literally free, but they are usually cheap, repeatable, and do not incur model-token spend. Stochastic operations add model latency and usage against a finite budget. That makes model calls the resource this architecture has to justify, especially when Part 5 tests whether extra sampling buys additional verified coverage.

Spending model budget on budget arithmetic, sentence extraction, or change detection is not merely inelegant. It spends the variable-cost resource on operations ordinary software can already perform exactly. The determinism default is therefore also a cost-control mechanism. This is why the runtime records usage and cost for every attempt (Chapters 11 and 13) and why its scheduler treats an exhausted budget as a first-class stopping condition (Chapter 28).

Do this now

Twenty minutes. Classify a real pipeline.

  1. Take one workflow you already run with AI — yours, or one at work. Write down every operation it performs, the way this chapter listed seven for the paragraph review. Aim for at least six.
  2. For each, run the test: can you write down what a correct answer would be, without writing down the answer? Mark it D (deterministic) or S (stochastic).
  3. Apply the annoyance test to every S: would you be irritated to get a different answer on a re-run? Every yes is misclassified. Move it.
  4. Count how many operations currently go through a model, and how many actually need to.

The gap between those two numbers is your entire cost-reduction program, and Chapter 6 puts a price on it.

Failure modes

  • Routing exact operations through the model. Asking for a token count, a date comparison, or a budget decision from a component that samples.
  • Treating temperature=0 as determinism. Measured false (He et al., 2025; Yuan et al., 2025). It removes the sampler and leaves the batch dependence.
  • A second model call to parse the first one’s output. Doubles the stochastic surface to avoid constraining an interface.
  • Averaging away disagreement. When two competent judgments differ, the spread is data. Collapsing it to a mean discards what the stochastic component actually produced.
  • Over-applying the rule. A rule cascade whose specification cost exceeds the value of determinism, maintained forever, failing silently.
  • Single-run belief. One passing result is a draw, not an estimate (Song et al., 2024).

What this chapter established

  • The governing rule is a strong default: if an operation can be computed from explicit state and rules, keep it deterministic rather than routing it through a model.
  • The operational test separates two questions: can you implement the operation deterministically, and if not, can you still check its result mechanically?
  • The productive asymmetry: where you cannot write the generator but can write an adequate check, the architecture is a stochastic proposal behind a deterministic verifier.
  • Stochasticity is not a defect to minimize to zero. It is useful where the process needs proposals ordinary software cannot economically enumerate, and it should be confined rather than allowed to govern the process.
  • temperature=0 removes sampling randomness, not every source of caller-visible variation. In Thinking Machines Lab’s controlled inference experiment, 1,000 identical greedy requests produced 80 distinct completions until batch-invariant kernels were substituted.
  • The override condition: when a deterministic specification is impractical relative to its value, use a model deliberately, preserve the boundary, and revisit it when the system changes.

Next

We now have a rule for which side of the line an operation belongs on, and an architecture for the cases where you cannot write the solution but can write the check: a stochastic generator behind a deterministic verifier.

That architecture is not hypothetical, and the next chapter goes looking for it in the wild. It turns out that one industry spent twenty years building exactly that verifier — compilers, type systems, test suites, continuous integration — for reasons that had nothing to do with AI, while simultaneously producing millions of written specifications and filing them under project management. That is why software automated itself first, and it is not because code is easy. The chapter turns the observation into a survey you can run on any domain.

Continue with Scrum Built the Training Set.

References