← Applied AI

Beyond the Chat Box

A chat window is an API whose integration layer is a person: in chat, you are the runtime. This chapter makes that hidden work visible, states the book's two bets, and takes the first step out of the chat box by moving one responsibility out of human memory and into the process.

Part 1 — Where You Stand

Before any of this can be built, it is worth being exact about where you are standing. Intelligence is no longer the scarce component. What is scarce is everything around it: deciding which work should never be stochastic, knowing what a model call costs, noticing that a person’s review decays as the model improves, and measuring a system whose output changes between runs.

These eight chapters make that case: the position a commodity technology erodes, the rule that keeps deterministic work deterministic, where this era’s training data came from, why human review becomes a rubber stamp, what intelligence actually costs, how to measure a stochastic system, and why projects still are not finishing.

The conclusion moves the target. Generation is not the bottleneck. Intent, authority and verification are.


The missing operation

CodeAI has a call whose shape is almost disappointingly ordinary:

result = adapter.prompt(paragraph)

It sends one request through the OpenCode gateway and returns a CallResult. It does not open a conversation window. It does not know which chapter you are writing. It does not decide whether the returned paragraph belongs in the manuscript.

That small separation is where this book begins.

Here is the work it separates. You are reviewing a paragraph for unsupported factual claims. You copy the paragraph into a chat window, explain what counts as support, and ask for a review. The model flags a sentence that needs a source. You find the source, decide whether it actually supports the sentence, and either add a citation or rewrite the sentence.

That is a useful process. It is also much larger than the two messages visible in the window.

You chose the paragraph and remembered which manuscript it came from. You supplied the review criteria, distinguishing a suggestion from a correction. You checked the source, decided the work was finished, and moved the result back into the right file.

The model participated. You operated every connection around it.

Draw the two arrangements side by side and the whole book is visible in outline:

CHAT

person
 ├─ remembers the objective
 ├─ selects the context
 ├─ calls the model
 ├─ interprets the result
 ├─ checks it
 ├─ moves it
 └─ decides what happens next

APPLIED AI

person supplies intent
        ↓
process holds state, context and rules
        ↓
model supplies a proposal
        ↓
process preserves, checks and routes it
        ↓
person retains authority

In the first arrangement the chat looks simple because you are the runtime. In the second, the model does the same job it did before; what changed is where everything around it lives. If you forget every class name in this book, remember that picture.

What changes when AI stops being something you converse with and becomes something your software can call?

The answer is not “it gets faster.” The answer is that every one of those connections becomes an engineering responsibility with a name, a failure mode, and a place to put a test.

The chat box is an API with a person in the middle

A text box is a human-operated API. The input field accepts a request, the output field returns a result, and a person translates between that exchange and the application that actually needs the work done.

There is nothing disgraceful about this arrangement. For a single question, a person is the best integration layer. They resolve ambiguity without a schema and notice when an answer is beside the point.

But run the paragraph review across a whole manuscript and the hidden work surfaces as questions nobody can answer. Which paragraphs have been reviewed? Which version did the model see? Did the last answer refer to the current paragraph or the one before the edit? Was the revised sentence saved, or is it still in a scroll buffer? If the application closes, where does the work resume?

Copy and paste moves bytes. It leaves task identity and state in the operator’s head, where it cannot be queried, tested, or resumed.

That is the chapter’s central observation, and it gives Applied AI its first method. Find the work the person is invisibly doing around the chat box, and move the repeatable parts of it into software. Not all of it at once, and not the judgment. The remembering, the matching of answers to inputs, the record of what was decided — the parts that are the same every time and fail silently when a tired person does them.

The obvious response is to get better at prompting. That response has been studied, and it does not scale the way people expect. Zamfirescu-Pereira and colleagues gave ten participants without significant prompt-design experience a no-code tool for building an instructional chatbot on GPT-3, then watched how they worked.

The participants explored prompt designs opportunistically rather than systematically: they overgeneralized from a single success, treated one good output as evidence that a prompt was robust, and wrote prompts shaped by expectations imported from instructing another human (Zamfirescu-Pereira et al., 2023). The authors connect these struggles to the long-documented difficulties of end-user programming and interactive machine learning.

Bound that result honestly. Ten participants, one task, one model generation, non-experts by design. It does not show that skilled engineers cannot write good prompts. The study shows how these non-experts approached prompt design in that interface. The architectural inference this book draws is separate: a text box by itself provides no explicit place to record evaluation criteria, controls, or regression tests.

Structure beat a better prompt

The complementary result is the more interesting one, because it changed the artifact rather than the operator.

Wu, Terry, and Cai took complex tasks that a single large prompt handled poorly and decomposed them into chains: sequences of small LLM operations where the output of one step becomes the input of the next, with every intermediate result exposed and editable. In a 20-person study, chaining improved the quality of task outcomes and significantly increased transparency, controllability, and users’ sense of collaboration (Wu et al., 2022).

The behaviors participants developed matter more here than the quality delta. They used sub-tasks to calibrate what they could expect from the model, compared strategies by watching parallel downstream effects, and debugged unexpected output by “unit-testing” individual sub-components of a chain.

That last observation is the one this book uses. When intermediate results are inspectable, failures can be localized to a step instead of disappearing inside one prompt-response pair. The model did not change; the structure around the call did.

That is this book’s thesis in an early, modest form: structure creates places to inspect, compare, and check model work.

Bound this one too: 20 participants, an interactive research prototype, GPT-3-era models, self-reported transparency and control alongside task quality. It demonstrates that decomposition gave these participants leverage. It does not prove chains are more accurate than monolithic prompts in general, and a badly-cut chain can be worse than one good prompt.

The first boundary

So we stop optimizing the text box and make the operation callable. At the client boundary, the process narrows to this:

explicit prompt → HTTP request → provider payload → CallResult

Here is a constructed teaching example using CodeAI’s actual adapter. The paragraph and the review task are illustrative; this is not a captured production run.

from codeai import OpenCodeCognitionAdapter

paragraph = "The service processes every request within one second."

adapter = OpenCodeCognitionAdapter(model="mimo-v2.5", protocol="chat_completions")
result = adapter.prompt(
    paragraph,
    instruction=(
        "Identify factual claims that require evidence. "
        "Do not invent sources. Return a short review."
    ),
)

if result.status != "succeeded":
    raise RuntimeError(result.error)
review = result.raw_output

That is the whole call: a model, a prompt, an instruction, and a result whose status says whether the call succeeded at this boundary. It assumes the CodeAI checkout is installed and an OPENCODE_ZEN_API_KEY is configured. The model identifier and its chat-completions route come from CodeAI’s inspected route table, not a promise of gateway availability.

prompt() is a convenience. Underneath it builds the identifiers, the context package, and the CallSpec that invoke() actually takes — inventing the identifiers because you did not supply any. Chapter 11 opens that up and makes each piece a recorded fact, without requiring a live service.

Two things changed. The program now holds a value named review that it can store, hash, or pass onward. And failures arrive through status and error rather than being returned as raw_output, so a failed call does not have to masquerade as review content.

One thing did not change: nothing has been verified. A generated review is a proposal about where to look. No amount of transport correctness establishes whether the service actually meets its latency promise.

From memory to fields — a constructed teaching example

The paragraph review named seven operator jobs: choosing, remembering, supplying criteria, distinguishing, checking, deciding, moving. Start with three of them. Each becomes one stored field; the record below is illustrative, not a captured run:

Operator jobStored fieldFailure it catches
Chose the paragraphparagraph_id, paragraph_shaReview attached to an edited paragraph
Supplied the criteriaprompt_versionTwo reviews judged by different rules
Moved the result backdisposition: pending / filedWork finished twice or never
import hashlib

paragraph_sha = hashlib.sha256(paragraph.encode()).hexdigest()
record = {
    "paragraph_id": "ch3-para-014",
    "paragraph_sha": paragraph_sha,
    "prompt_version": "claims-review-v1",
    "model": "mimo-v2.5",
    "review": review,          # content channel only; never an error string
    "disposition": "pending",
}

You can now do one new thing: hash any paragraph, attach the review to that hash rather than to memory, and refuse to file a review whose paragraph_sha no longer matches the file. That is exact input binding in its smallest form. Chapter 11 will add the separate task, call, and attempt identities that let the process say which work this observation belongs to.

These fields are not more architecture. They are the first three pieces of the operator’s memory leaving the chat box: which paragraph was this about, which rules judged it, and has it been dealt with. Before, only you knew. Now the program does.

The adapter makes this boundary unusually easy to see. For the chat_completions route used above, prepare() builds a one-message request and invoke() sends the prepared request. The returned CallResult carries the output, usage provenance through usage_source, and failures through status and error. prompt() adds no retries, conversation history, or ledger.

That is a source-level observation about one inspected adapter. It does not imply the remote service stores nothing, or that the rest of CodeAI is stateless. The point is ownership: this object does not own the review process. Which forces the useful question — what does?

Intelligence inside a process

The working definition for this book:

A model gives you intelligence. Applied AI is the engineering required to make that intelligence participate reliably in a process.

By the end of the book the same idea has a sharper form: the model is not the process. It is one selectively invoked component inside a process you can inspect.

“Reliably” needs a scope, or it is marketing. It does not mean every model answer is correct. It means the application can distinguish outcomes, retain what happened, control effects, and obtain the evidence its next decision requires — including reporting honestly that a task ended unresolved.

Here is where the book ends up:

    flowchart TD
    S["software state"] --> C["context assembly<br/><i>name exactly what the model sees</i>"]
    C --> M["model call<br/><i>the only stochastic step</i>"]
    M --> R["raw result<br/><i>preserved before interpretation</i>"]
    R --> E["evaluation<br/><i>what does this support?</i>"]
    E --> D["decision<br/><i>recorded against claims</i>"]
    D --> A["action<br/><i>authorized, bounded</i>"]
    A --> V["verification<br/><i>independent observation</i>"]
    V --> S2["new software state"]
    S2 -.-> S
  

Don’t learn this diagram yet. At this point it means only one thing: everything the person was doing implicitly needs somewhere explicit to live. Most of the boxes are jobs from the chat arrangement above, given a home. Chapters 11 through 28 earn the boxes one at a time, and not every arrow fires on every pass.

The two bets

Most of this book is ordinary engineering. Two claims are not, and it is fairer to state them now than to spring them later. Later chapters have to earn them.

Bet one: exactly one box in that diagram is allowed to be stochastic. Context assembly, preservation, evaluation against recorded mechanical criteria, authorization, state transition, and mechanical verification belong on the deterministic side wherever their specifications can be written. Only the model call is permitted to generate from a space you could not have enumerated in advance. The aim is not to shrink that box to nothing — its ability to propose what you did not anticipate is what you are paying for — but to confine it, with a record of what went in and what came out and independent evidence on the other side. Chapter 3 argues the rule properly.

Bet two: the thing that decides what operation the process needs next must not itself be intelligent. The book builds toward a small deterministic scheduler that chooses the next operation from explicit state and policy. Only after that decision is CALL does model selection become relevant; choosing an occupant is a separate problem from choosing whether another model call is needed at all. A stochastic governor over a stochastic component adds a layer to the reliability problem and makes failures harder to locate. The book also makes a stronger bet, that the governing policy could be a readable specification rather than code. Chapter 28 reports what became of both.

One project, one construction story

The project you will build is CodeAI, with OpenCode as the primary model gateway for the walkthroughs. Cost is part of the reason: if the process is going to make many small calls, the per-call price is an architectural constraint, not an accounting detail. Gateway pricing and route availability are verified during live labs rather than advertised here.

The construction starts exactly where this chapter did — one callable model operation — and builds the process around it, one responsibility at a time. You will get prompts for each step, not only finished diagrams.

The pattern is constant: inspect the existing project, request one bounded increment, run it, check the observed result, keep the evidence. The prompts work with Codex and Claude as well; the engineering obligations do not change with the assistant. Where this book describes a stronger contract than the code implements, it says so.

CodeAI is the starting point for a reason. It was not bought; it grew, one missing capability at a time, around one author’s way of working. Building software around one person’s workflow has historically been expensive; as construction gets cheaper, it becomes increasingly feasible. The chat window will not disappear from your day, and it should not. What should leave it is the process, and Chapter 30 argues that the process this book builds is what makes tools shaped around your own work inspectable and governable.

The first step out of the chat box

Go back to what this chapter actually built. A review that used to float in a scroll buffer is now a value attached to the hash of the paragraph it judged, the version of the rules that judged it, and a disposition that says whether it has been dealt with.

We have not built an agent. We have not verified anything, and the model is no smarter than it was. We have moved one responsibility out of human memory and into the process. That is Applied AI at its smallest size, and every later chapter makes the same move for another responsibility: what the model saw, what came back, what was claimed, who may act, and whether the work is done.

Failure modes

  • Reading an error as content. A diagnostic string assigned to review and filed in the manuscript. Solved at the client boundary in Chapter 11.
  • Associating a result with the wrong input. A delayed response attached to an edited paragraph. Needs task identity (Chapter 11) and an explicit context package (Chapter 15).
  • Treating a suggestion as an authorized action. “Add a citation here” silently becoming a file write. Needs the capability/authority split (Chapter 20).
  • Mistaking a successful request for a completed task. HTTP 200 is a transport fact, not a work outcome.
  • Generalizing from one good output. The study above documents that behavior among non-experts. The engineering rule is broader because a stochastic system cannot be characterized from one draw: one passing run is a sample, not an estimate.

What this chapter established

  • In a chat workflow the person is the runtime: they route context, hold state, judge output, and transfer effects. Calling the model directly makes those responsibilities available to software; it does not discharge them. Applied AI begins by moving the repeatable ones into the process.
  • Zamfirescu-Pereira and colleagues document how non-experts struggled to design prompts systematically; the book’s architectural inference is that a text box alone provides no explicit structure for controls, checks, or regression tests (Zamfirescu-Pereira et al., 2023).
  • In the AI Chains study, decomposing work into inspectable steps improved task outcomes, and participants debugged unexpected outputs by “unit-testing” sub-components of a Chain (Wu et al., 2022).
  • The definition we will hold to scope: a dependable process distinguishes outcomes, retains what happened, controls effects, and can report an unresolved result honestly.
  • Two bets, stated here and earned later: only one component may be stochastic (Chapter 3), and the router that governs it must not be (Chapter 28).
  • Obtaining a review value would demonstrate callable generation. It would not demonstrate factual verification, safe editing, restartability, or useful model selection. No live model outcome is claimed in this chapter.
  • The chapter’s one step: a review attached to a paragraph hash, a prompt version and a disposition. No agent and no verification, but one responsibility moved out of human memory and into the process.

Next

Before any of this is worth building, there is a prior question: why should you be the one building it? That diagram assigns the human intent, authority, and verification — plus judgment about when the model is the right tool at all. The next chapter argues that those four are the only positions in the picture worth holding, states plainly that this book will be wrong about some things, and shows what being wrong responsibly looks like using a study that reversed itself.

Continue with Never Stand in Front of the Steamroller.

References

  • J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ‘23), Article 437, 21 pages. https://doi.org/10.1145/3544548.3581388
  • Tongshuang Wu, Michael Terry, and Carrie Jun Cai. AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ‘22), Article 385, 22 pages. https://doi.org/10.1145/3491102.3517582

Implementation source: CodeAI, src/codeai/providers.py (OpenCodeCognitionAdapter.prompt, OpenCodeCognitionAdapter.invoke), src/codeai/domain.py (CallSpec, ActorRef), src/codeai/context.py (ContextCompiler.compile). Source identities and limitations are recorded in docs/applied-ai/evidence-map.md in the book repository.