<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Applied AI on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/books/applied-ai/</link><description>Recent content in Applied AI on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Mon, 14 Sep 2026 05:00:30 +0100</lastBuildDate><atom:link href="https://aibussin.com/books/applied-ai/index.xml" rel="self" type="application/rss+xml"/><item><title>Beyond the Chat Box</title><link>https://aibussin.com/books/applied-ai/01-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:01 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/01-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Before any of this can be built, it is worth being exact about where you are standing. Intelligence is no longer the scarce component. What is scarce is everything around it: deciding which work should never be stochastic, knowing what a model call costs, noticing that a person&amp;rsquo;s review decays as the model improves, and measuring a system whose output changes between runs.&lt;/p&gt;&#10;&lt;p&gt;These eight chapters make that case: the position a commodity technology erodes, the rule that keeps deterministic work deterministic, where this era&amp;rsquo;s training data came from, why human review becomes a rubber stamp, what intelligence actually costs, how to measure a stochastic system, and why projects still are not finishing.&lt;/p&gt;</description></item><item><title>Never Stand in Front of the Steamroller</title><link>https://aibussin.com/books/applied-ai/02-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:02 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/02-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-plain-version"&gt;The plain version&lt;/h2&gt;&#10;&lt;p&gt;The historical premise of this chapter is not controversial: machines have displaced codified parts of skilled work before. The inference for AI — which parts are now exposed, how fast the boundary moves, and where durable human responsibility remains — is the claim this chapter has to earn.&lt;/p&gt;&#10;&lt;p&gt;&amp;ldquo;Computer&amp;rdquo; was a job title before it was a machine: rooms of people did calculations by hand for engineers, observatories and space programs. Typesetters set newspapers line by line, first by hand and later on hot-metal machines, until photocomposition and then desktop publishing took the trade. Keypunch operators turned programs and data into stacks of punched cards. Typing pools produced every letter an office sent. Each of those was skilled work, done by people who were good at it and some who were superb. None of that skill protected the job once a machine could do the codified part.&lt;/p&gt;</description></item><item><title>If There's Any Doubt, It's Deterministic</title><link>https://aibussin.com/books/applied-ai/03-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:03 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/03-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="seven-operations-one-of-them-intelligent"&gt;Seven operations, one of them intelligent&lt;/h2&gt;&#10;&lt;p&gt;Take the paragraph review from Chapter 1 and write down everything the process actually has to do. Not the prompt — the process.&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Select the paragraph to review.&lt;/li&gt;&#10;&lt;li&gt;Decide whether it has changed since the last review.&lt;/li&gt;&#10;&lt;li&gt;Decide whether there is budget left for a call.&lt;/li&gt;&#10;&lt;li&gt;Identify claims in the paragraph that need supporting evidence.&lt;/li&gt;&#10;&lt;li&gt;Extract the flagged sentence and its position.&lt;/li&gt;&#10;&lt;li&gt;Check that a cited source resolves and that its year matches the bibliography.&lt;/li&gt;&#10;&lt;li&gt;Decide whether to write the change to the file.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Six of those have exactly one correct answer, computable without a model. Selection is an index lookup and change detection is a hash comparison. Budget is arithmetic; extraction is parsing. Source checking resolves the request and compares it against the bibliography. The write decision evaluates policy.&lt;/p&gt;</description></item><item><title>Scrum Built the Training Set</title><link>https://aibussin.com/books/applied-ai/04-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:04 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/04-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Why checkable work became unusually favorable terrain&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="an-uncomfortable-piece-of-bookkeeping"&gt;An uncomfortable piece of bookkeeping&lt;/h2&gt;&#10;&lt;p&gt;Software engineering became an unusually fertile early target for language-model automation. The usual explanation is that code is logical and therefore tractable for a machine.&lt;/p&gt;&#10;&lt;p&gt;That explanation points at only part of what made software favorable. The account this chapter defends is about the surrounding work records and checks.&lt;/p&gt;&#10;&lt;p&gt;Over decades, software teams increasingly worked through issue trackers, version control, code review, automated tests, and iterative planning methods. Not every team used Scrum, and none of these practices guarantees the others, but together they often left behind unusually structured records:&lt;/p&gt;</description></item><item><title>Meat Proxy</title><link>https://aibussin.com/books/applied-ai/05-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:05 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/05-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Review degrades because the system succeeds&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="ten-weeks"&gt;Ten weeks&lt;/h2&gt;&#10;&lt;p&gt;Week one, the model drafts a paragraph review and you read every line of it. You disagree with two of its four flags, check a source yourself, rewrite one sentence. The whole thing takes twenty minutes and you feel you have earned the output.&lt;/p&gt;&#10;&lt;p&gt;Week three, the reviews have been good. You read them properly but you stop re-checking the sources it says it checked. Nothing bad happens.&lt;/p&gt;</description></item><item><title>The Price of Intelligence</title><link>https://aibussin.com/books/applied-ai/06-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:06 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/06-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-invoice-nobody-runs-at-prototype-time"&gt;The invoice nobody runs at prototype time&lt;/h2&gt;&#10;&lt;p&gt;The paragraph review from Chapter 1 costs about 1,200 tokens in and 400 out. At prototype scale that is a rounding error, and it is why nobody computes the next four numbers.&lt;/p&gt;&#10;&lt;p&gt;A manuscript has 3,000 paragraphs. That is one full pass.&lt;/p&gt;&#10;&lt;p&gt;The review is not run once. Every edit invalidates the review of the paragraph it touched, and a careful author edits most paragraphs four or five times.&lt;/p&gt;</description></item><item><title>Intelligence in the Wrong Direction</title><link>https://aibussin.com/books/applied-ai/07-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:07 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/07-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Capability without an objective is magnitude without direction&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="the-thing-the-marketing-does-not-cover"&gt;The thing the marketing does not cover&lt;/h2&gt;&#10;&lt;p&gt;Every vendor will sell you more capability. None of them can tell you whether that capability moves your process toward its objective without a measurement you supply.&lt;/p&gt;&#10;&lt;p&gt;A stronger model can execute the wrong objective more effectively just as it can execute the right one more effectively. Capability alone does not tell you whether the gap between what you wanted and what you got narrowed or widened. That is an empirical question.&lt;/p&gt;</description></item><item><title>Where Are the Finished Projects?</title><link>https://aibussin.com/books/applied-ai/08-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:08 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/08-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-tireless-coworker"&gt;The tireless coworker&lt;/h2&gt;&#10;&lt;p&gt;Imagine you are given a colleague with the properties the industry has been describing for nearly four years.&lt;/p&gt;&#10;&lt;p&gt;They never tire. They work through the night and the weekend. They have read essentially everything and can recall it instantly. They write code faster than you can read it, and they will take on the tedious parts you have been avoiding since March.&lt;/p&gt;</description></item><item><title>One Runtime, Many Windows</title><link>https://aibussin.com/books/applied-ai/09-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:09 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/09-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;A chat window is an interface with a person inside the loop doing the integration. This part takes the process out of it and gives the work somewhere else to live.&lt;/p&gt;&#10;&lt;p&gt;That happens in stages. The work needs a durable identity and a record that survives the interface closing. The model underneath has to be replaceable, because it will be replaced. A model call has to become a recorded process event rather than a string. That call then has to survive a change of provider dialect, its usage numbers have to mean something before anyone does arithmetic on them, and finally a successful call has to stop being mistaken for finished work.&lt;/p&gt;</description></item><item><title>A Revolver, Not a Foundation</title><link>https://aibussin.com/books/applied-ai/10-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:10 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/10-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;The model will change; the frame should survive it&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="the-book-you-could-write-better-next-year"&gt;The book you could write better next year&lt;/h2&gt;&#10;&lt;p&gt;Write a book with AI this year and you will be able to write the same book better next year. And better again the year after. The same is true of the code, the research synthesis, the review process, the classifier — anything where a model does part of the work.&lt;/p&gt;</description></item><item><title>The Smallest Useful Model Call</title><link>https://aibussin.com/books/applied-ai/11-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:11 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/11-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-smallest-working-call"&gt;The smallest working call&lt;/h2&gt;&#10;&lt;p&gt;CodeAI&amp;rsquo;s OpenCode cognition adapter, invoked directly without the recorded-call runtime, does the minimum well. Chapter 1 reached it through &lt;code&gt;prompt()&lt;/code&gt;, which builds the &lt;code&gt;CallSpec&lt;/code&gt; for you. Underneath, one &lt;code&gt;CallSpec&lt;/code&gt; becomes one user message, the message goes out as an ordinary HTTP request, and a &lt;code&gt;CallResult&lt;/code&gt; comes back:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;adapter &lt;span style="color:#f92672"&gt;=&lt;/span&gt; OpenCodeCognitionAdapter(model&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;mimo-v2.5&amp;#34;&lt;/span&gt;, protocol&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#e6db74"&gt;&amp;#34;chat_completions&amp;#34;&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;result &lt;span style="color:#f92672"&gt;=&lt;/span&gt; adapter&lt;span style="color:#f92672"&gt;.&lt;/span&gt;invoke(spec) &lt;span style="color:#75715e"&gt;# one CallSpec in, one CallResult back&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;if&lt;/span&gt; result&lt;span style="color:#f92672"&gt;.&lt;/span&gt;status &lt;span style="color:#f92672"&gt;!=&lt;/span&gt; &lt;span style="color:#e6db74"&gt;&amp;#34;succeeded&amp;#34;&lt;/span&gt;:&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;raise&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;RuntimeError&lt;/span&gt;(result&lt;span style="color:#f92672"&gt;.&lt;/span&gt;error)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;review &lt;span style="color:#f92672"&gt;=&lt;/span&gt; result&lt;span style="color:#f92672"&gt;.&lt;/span&gt;raw_output&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It keeps failure in its own channel instead of passing a diagnostic off as a review. It returns the usage with its source label alongside the convenient string. Those two habits already put it ahead of a great deal of production code.&lt;/p&gt;</description></item><item><title>One Operation, Several Model APIs</title><link>https://aibussin.com/books/applied-ai/12-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:12 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/12-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Choosing a model can change the whole wire contract&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="right-model-wrong-dialect"&gt;Right model, wrong dialect&lt;/h2&gt;&#10;&lt;p&gt;Chapter 11&amp;rsquo;s second live run failed with an HTTP 500. The gateway, the model, and the credential were all right. The request was shaped for OpenCode&amp;rsquo;s Responses endpoint, and &lt;code&gt;mimo-v2.5&lt;/code&gt; is served on Chat Completions.&lt;/p&gt;&#10;&lt;p&gt;That looks like a configuration slip. It is a property of the ground you are building on.&lt;/p&gt;</description></item><item><title>Token Counts Don't Add Up</title><link>https://aibussin.com/books/applied-ai/13-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:13 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/13-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;A field name is not a unit: normalize at the boundary&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Every AI application ends up doing arithmetic on numbers a provider reported. It adds tokens into a budget, divides an invoice by them, compares two models on them, and routes work to whichever looks cheaper per token.&lt;/p&gt;&#10;&lt;p&gt;Those numbers arrive in fields with shared names: &lt;code&gt;input_tokens&lt;/code&gt;, &lt;code&gt;cached_tokens&lt;/code&gt;, &lt;code&gt;reasoning_tokens&lt;/code&gt;. The shared names hide different rules. On one major API, cached input is counted &lt;em&gt;inside&lt;/em&gt; &lt;code&gt;input_tokens&lt;/code&gt;. On another, it is reported &lt;em&gt;beside&lt;/em&gt; it. Add the two and the result is wrong even though every number in it is correct. Nothing crashes. The budget, the cost comparison and the routing decision are simply built on a sum that means nothing.&lt;/p&gt;</description></item><item><title>A Successful Call Is Not Finished Work</title><link>https://aibussin.com/books/applied-ai/14-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:14 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/14-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;The model is not the process&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;An AI pipeline gets a response back, the response looks right, and something marks the work done. That step is where many AI-enabled processes quietly lose track of the truth.&lt;/p&gt;&#10;&lt;p&gt;A successful model call tells you the call worked. It does not tell you that anyone checked the output against what the task required, that the thing checked is the thing being used, that someone with standing agreed, or that the agreement was written down. Collapse those into one status field and &amp;ldquo;done&amp;rdquo; starts to mean &amp;ldquo;the model returned something plausible&amp;rdquo;.&lt;/p&gt;</description></item><item><title>What Did the Model Actually See?</title><link>https://aibussin.com/books/applied-ai/15-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:15 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/15-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 3 — Give Intelligence a Runtime&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;A component with a contract still needs a world around it. This part builds that world out of the things that used to live in a person&amp;rsquo;s head and in the scrollback.&lt;/p&gt;&#10;&lt;p&gt;Each chapter makes one of them explicit and durable. What the model was allowed to see becomes a compiled, recorded input rather than an accumulated transcript. Working state becomes something another process can resume from instead of restarting blind. The provider&amp;rsquo;s response is preserved before anything interprets it, so a corrected reading can be applied later without asking the model again. Statements become claims that point at evidence, and decisions record what they relied on. Then the process is allowed to change something outside itself — and to observe the result rather than trust a report of it.&lt;/p&gt;</description></item><item><title>Restart Is Not Resume</title><link>https://aibussin.com/books/applied-ai/16-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:16 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/16-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 3 — Give Intelligence a Runtime&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Externalize working memory&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;When an AI process dies partway through, the usual fix is to restart it and re-run whatever did not finish. For a model call, that can mean paying for the same request twice. For an agent that sends email, opens tickets, charges a card or edits files, it can mean doing something twice that cannot be undone.&lt;/p&gt;&#10;&lt;p&gt;The capability you actually need is a different one:&lt;/p&gt;</description></item><item><title>Preserve Before You Interpret</title><link>https://aibussin.com/books/applied-ai/17-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:17 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/17-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 3 — Give Intelligence a Runtime&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Raw output first&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Every rule an AI application applies to a model&amp;rsquo;s response is an interpretation. Did generation finish? Is this valid JSON? Is this a refusal? How many tokens did it cost? What does the answer actually say? Those rules can turn out to be wrong: a provider adds a field, a parser misses a case, a failure shows up that nobody anticipated.&lt;/p&gt;</description></item><item><title>Claims, Evidence, and Decisions</title><link>https://aibussin.com/books/applied-ai/18-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:18 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/18-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 3 — Give Intelligence a Runtime&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;What did the decision rest on?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;AI output feeds decisions. A review says a setting is correct, a summary says the tests pass, an assistant says a change is safe, and someone merges. When one of those statements turns out to be wrong, asking whether the AI was wrong gets you very little. What you need to know is which statements the decision actually rested on, and what anyone had checked. Most systems cannot tell you, because the decision was recorded without its basis.&lt;/p&gt;</description></item><item><title>The Agent Said Done. Did Anything Change?</title><link>https://aibussin.com/books/applied-ai/19-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:19 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/19-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 3 — Give Intelligence a Runtime&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Observe the effect, not the report&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Sooner or later an AI application is allowed to change something: edit a file, update a record, send a message, merge a change. The moment it does, a new kind of mistake becomes possible. A worker reports “done”, the report is recorded as the outcome, and nobody looks at what actually changed. The report can be sincere and still wrong. The edit went to the wrong file, stopped halfway, or never happened.&lt;/p&gt;</description></item><item><title>Capability Is Not Authority</title><link>https://aibussin.com/books/applied-ai/20-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:20 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/20-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 4 — Make It Safe and Verifiable&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;The process can now act on the world. That is exactly when the safety questions arrive, and they are not questions a prompt can answer.&lt;/p&gt;&#10;&lt;p&gt;Three of them run through this part. Being able to perform an action is not being permitted to perform it, so permission has to be checked in code, outside the text a model generates. A check that passes is not proof, so verification has to be independent of the generator, adequate to the property that matters, and bound to the exact state being accepted. And trying again is not free, because a retry can cause a second real effect where a replay would have returned a record.&lt;/p&gt;</description></item><item><title>The Agent Cannot Grade Its Own Homework</title><link>https://aibussin.com/books/applied-ai/21-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:21 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/21-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 4 — Make It Safe and Verifiable&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Independence, adequacy, and binding&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;A model saying its work is correct is not verification. Neither, on its own, is running a separate test. A test can run separately and check the wrong thing. It can check the right thing against a version of the artifact that has since changed. And a PASS that nobody can tie to the exact state being accepted is only another claim.&lt;/p&gt;</description></item><item><title>Retries Are Side Effects Too</title><link>https://aibussin.com/books/applied-ai/22-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:22 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/22-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 4 — Make It Safe and Verifiable&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Retry, replay, duplicate&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Retrying is the most ordinary recovery move in software, and in an AI application it is often the most dangerous one. When a call or an action fails, or only seems to, the loop tries again. If the first attempt already sent the email, charged the card, wrote the file or ran the model, the retry does it twice. A timeout says nothing about whether the far side acted.&lt;/p&gt;</description></item><item><title>Blind Before You Compare</title><link>https://aibussin.com/books/applied-ai/23-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:23 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/23-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 5 — More Intelligence Is Not Automatically Better&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Once a system can produce several answers, adding more intelligence looks easy: call the model again, try another model, change the prompt, build a portfolio. None of that guarantees more useful coverage.&lt;/p&gt;&#10;&lt;p&gt;This part builds the comparison in stages. Proposals first have to be collected without influencing one another, because agreement after exposure means something different from agreement without it. Then variety has to be measured by what actually gets solved rather than by how many model names were involved. When the first benchmark turns out to be too easy to tell the methods apart, the test itself has to be made more discriminating. A promising signal found inside a failed experiment then has to survive a matched replication before it is allowed to change anything.&lt;/p&gt;</description></item><item><title>The Models Were Different. Their Mistakes Weren't</title><link>https://aibussin.com/books/applied-ai/24-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:24 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/24-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 5 — More Intelligence Is Not Automatically Better&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Measure variety by what gets solved&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Reach for a second model and the reasoning feels obvious. Different training, different weaknesses, so where one fails another may succeed. Add a third and you have a portfolio. It is an appealing story, and it is a hypothesis rather than a property of model names.&lt;/p&gt;&#10;&lt;p&gt;What you actually need is &lt;em&gt;complementary errors&lt;/em&gt;: cases where one generator fails and another succeeds. Different names do not guarantee that. Models trained on overlapping data, prompted the same way, on the same task, can be wrong in the same way. When they are, a portfolio costs more and covers nothing extra.&lt;/p&gt;</description></item><item><title>When Your Benchmark Is Too Easy</title><link>https://aibussin.com/books/applied-ai/25-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:25 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/25-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 5 — More Intelligence Is Not Automatically Better&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Make the problems harder&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Suppose you compare two ways of using AI on your own evaluation set, and both solve eleven of twelve tasks.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Method A 11/12&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Method B 11/12&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The obvious reading is that the methods are equally good. There is another reading. If eleven of those tasks are easy enough that any reasonable method solves them, the comparison had almost nowhere to show a difference. Two methods can only disagree on tasks where at least one of them could fail. When nearly everything passes, the evaluation loses its power to discriminate, and &amp;ldquo;no difference observed&amp;rdquo; stops being evidence of equivalence. It may only be evidence of a ceiling.&lt;/p&gt;</description></item><item><title>Discovery Is Not Promotion</title><link>https://aibussin.com/books/applied-ai/26-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:26 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/26-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 5 — More Intelligence Is Not Automatically Better&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Diversity without more models&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="the-subgroup-that-almost-rewrites-the-chapter"&gt;The subgroup that almost rewrites the chapter&lt;/h2&gt;&#10;&lt;p&gt;Start with the temptation, because the discipline only means something if the temptation is real. Inside a failed experiment sits this: one wording — counterfactual — with three draws per task covered 9 of 12 tasks. The normal wording, given three draws per task in the same portfolio, covered 7. Read quickly, that is a better prompt.&lt;/p&gt;</description></item><item><title>Replicate Before You Believe</title><link>https://aibussin.com/books/applied-ai/27-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:27 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/27-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 5 — More Intelligence Is Not Automatically Better&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;A signal is not a result, and a result is not a default&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Somewhere in your own records there is an encouraging pattern: a prompt that seemed to work better, a model that looked stronger on the cases you happened to inspect, a setting that went with success more often than not. The temptation is to adopt it.&lt;/p&gt;&#10;&lt;p&gt;The problem is where the pattern came from. It was found by looking at data you already had, which means the data chose the hypothesis. Anything strong enough to notice in a small sample is also the kind of thing chance produces regularly, and hindsight makes it feel predicted rather than discovered.&lt;/p&gt;</description></item><item><title>What Should Happen Next?</title><link>https://aibussin.com/books/applied-ai/28-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:28 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/28-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 6 — Put Intelligence Into the Process&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;The pieces are all built. This part puts them under one policy and then joins them.&lt;/p&gt;&#10;&lt;p&gt;It starts with the decision that usually gets skipped: which operation does the process need next — a model call, a deterministic check, a person, or a stop? That choice comes before any question about which model, and this part measures what climbing from cheap to expensive actually bought. Then one task runs through every boundary the book has built, from intent and authority to effect, verification and acceptance, to find out whether mechanisms that each work alone still hold where they hand off to one another.&lt;/p&gt;</description></item><item><title>One Process, End to End</title><link>https://aibussin.com/books/applied-ai/29-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:29 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/29-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 6 — Put Intelligence Into the Process&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Correct parts can still fail where they meet&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Every mechanism in this book was tested on its own, against a fixture built for its chapter. Context was compiled and recorded. Effects were observed by the runtime rather than taken from the worker&amp;rsquo;s report — or so each component&amp;rsquo;s own test reported. Checks were bound to the state they examined. Acceptance cited its evidence. Replays refused to act twice. Each passed its own test.&lt;/p&gt;</description></item><item><title>Your Applied AI</title><link>https://aibussin.com/books/applied-ai/30-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:30 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/30-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 6 — Put Intelligence Into the Process&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="you-will-still-open-the-chat-window-tomorrow"&gt;You will still open the chat window tomorrow&lt;/h2&gt;&#10;&lt;p&gt;You have read a book about getting AI out of the chat box. Tomorrow morning you will type something into one.&lt;/p&gt;&#10;&lt;p&gt;That is not a contradiction, and it is worth saying why before anything else. Conversation is the most direct way a person has to say what they want. Typing or talking to a model will stay part of how you work for as long as intent has to come from somewhere, and intent comes from you.&lt;/p&gt;</description></item></channel></rss>