Scrum Built the Training Set
Software became an unusually tractable early domain for AI because its work was already decomposed, versioned, checked, and recorded. Decades of software practice left behind task descriptions, tests, changes, attempts, and verdicts that modern benchmarks and training environments can reuse. This chapter shows how to read any domain for that structure, and how to build applications that keep those books on themselves.
Part 1 — Where You Stand
Why checkable work became unusually favorable terrain
An uncomfortable piece of bookkeeping
Software engineering became an unusually fertile early target for language-model automation. The usual explanation is that code is logical and therefore tractable for a machine.
That explanation points at only part of what made software favorable. The account this chapter defends is about the surrounding work records and checks.
Over decades, software teams increasingly worked through issue trackers, version control, code review, automated tests, and iterative planning methods. Not every team used Scrum, and none of these practices guarantees the others, but together they often left behind unusually structured records:
- a bounded unit of work,
- a natural-language description of what was wanted,
- tests or acceptance criteria that could help judge the result,
- explicit workflow state,
- written discussion and review,
- and a versioned change linked back to the work.
Read that as a data schema:
| Workflow artifact | What it can provide |
|---|---|
| Ticket description | A candidate natural-language task specification |
| Acceptance criteria / tests | A candidate check |
| The merged diff that closed it | A reference change |
| Failing builds, rejected reviews, reopened tickets | Failed or rejected attempts, when preserved with the reason |
| Comments, reviews, state transitions | A trace of how the work evolved |
The useful claim is not that software spent twenty years producing perfect (specification, verifier, solution) triples. Most records are messier than that. It is that software produced these ingredients at enormous scale, often in machine-readable systems and sometimes in public.
Nobody did this to train a model. The point is that later evaluation and training systems could reuse records created for ordinary software work.
Which domains are ready to be automated, and how would you tell before you spent the money finding out?
Why the records mattered
Machine learning is not new, but the connection here needs care. In 1986 Rumelhart, Hinton and Williams described backpropagation as adjusting network weights to reduce error between produced and desired outputs (Rumelhart, Hinton & Williams, 1986). Modern language-model pretraining is not simply that supervised recipe applied to software tickets, and this chapter is not claiming that Scrum records were the hidden cause of language-model capability.
The useful parallel is narrower: learning and evaluation become easier when a system has examples, objectives, and feedback, while software engineering independently became unusually good at leaving behind tasks, changes, and machine-executable checks.
That discipline is older than the word agile. Larman and Basili trace iterative and incremental development through the X-15 program of the 1950s, NASA’s Project Mercury, and IBM’s Federal Systems Division work on the first US Trident submarine (Larman & Basili, 2003). The commercial vocabulary arrived later. A 1986 Harvard Business Review article on Japanese product development used the rugby metaphor scrum (Takeuchi & Nonaka, 1986), and Sutherland and Schwaber later developed Scrum as a software method.
Put the histories beside each other and the engineering opportunity becomes visible without turning the parallel into causation. Software accumulated bounded work items, versioned attempts, tests, reviews, and outcomes for its own reasons. Later AI systems could use records with that shape for benchmarks, post-training environments, verifier training, and evaluation.
That is the claim to keep. Much of the public code corpus owes nothing to Scrum by name, and many software records are not clean training examples. The advantage is the availability and checkability of the records, not the brand of the ceremony.
The receipt
This is not only an argument by analogy. You can look at how the standard benchmark for AI software engineering was built.
Jimenez and colleagues constructed SWE-bench by taking roughly 90,000 pull requests from 12 widely-used Python repositories and filtering them on four criteria: the PR resolves a GitHub issue, it contributes tests, the project installs, and the tests exhibit a fail-to-pass transition — failing before the change and passing after it. What survived was 2,294 task instances, each evaluated with FAIL_TO_PASS and PASS_TO_PASS tests (Jimenez et al., 2024).
Look at what each filter contributes. Resolves an issue supplies a natural-language task description. Contributes tests supplies candidate checks. Fail-to-pass establishes that at least one task-specific test distinguishes the pre-fix and post-fix states. The diff supplies the reference patch.
Nobody wrote those 2,294 task records for SWE-bench. Developers wrote issues, tests, and patches to do software work; the benchmark later curated them into an evaluation format.
What the receipt does not show is how much of any particular model’s coding ability came from records of this shape rather than from public code in general; that split is not published for the models this book uses. Nor does existing bookkeeping make a benchmark or training environment free: SWE-bench still had to filter, install, execute, and normalize the harvested work. The narrower claim is enough. When a domain already records tasks, attempts, and discriminating checks, the raw material for evaluation and learning is much cheaper to assemble.
The failures were data too
SWE-bench kept the successes. The discipline recorded far more than successes: the build that went red, the review that sent a change back, the commit that was reverted, the ticket that was reopened. Each of those is also a specification, a verifier and an attempt. The verdict simply says no.
Those records are not waste. SWE-Gym is a training environment built the same way as SWE-bench, from 2,438 tasks in 11 Python repositories. Its authors ran agents against the tests and kept the trajectories. To train a verifier — a model that predicts whether an attempt solved its task — they used 1,318 trajectories that passed and 1,318 that did not (Pan et al., 2025). Half the training signal was failure. A verifier shown only successes would have had nothing to learn the difference from.
That gives the rule for everything that follows in this book. An attempt tied to its task and a meaningful verdict is a labeled example, whichever way the verdict went. An unchecked attempt may still be useful as preserved history, but it cannot tell you whether the work succeeded. For evaluation and improvement, the valuable record is the attempt together with what it was trying to do and what the check actually established.
You probably will not fine-tune a model on your records, and you do not need to for this to pay. At the scale of one team or one person, the same records can become a frozen task set for evaluation (Chapter 7), durable history another process can resume from (Chapter 16), preserved output a later interpreter can revisit (Chapter 17), and evidence for deciding whether a changed process or routing policy deserves promotion (Chapters 27–28). “Training set” is one possible use of these books. It is not the only one.
Why code was favorable early terrain
Code’s formal structure helps, but the more reusable advantage is the surrounding verification machinery.
Software became unusually favorable because it has unusually cheap automated checks.
Consider what a software engineer can often use before asking a person to inspect the result: a compiler that rejects whole classes of invalid programs, a type checker that enforces declared constraints, tests that exercise specified behavior, continuous integration that runs those checks automatically, and version control that makes many failed attempts cheap to discard.
None of those is a general proof of correctness. A program can compile, type-check, and pass its tests while still being wrong. Their value is that they provide fast, repeatable evidence about declared properties.
Chapter 3 gave the productive architecture: where you cannot write the generator but can write an adequate check, put a stochastic proposal behind deterministic verification. Software entered the AI era with decades of investment in exactly those kinds of checks.
This reframes the question. The useful axis is not just difficulty. It is checkability.
That distinction has teeth. Some hard programming problems come with executable tests, while apparently simple administrative tasks can lack a cheap observable criterion for success. The easier task is not necessarily the easier one to automate reliably.
How much a cheap check can prune
AlphaCode makes the leverage visible. It generated vast pools of candidate programs, then filtered them against the example tests included in the problem statement. Filtering on those example tests removed more than 99% of the model’s samples. With 10⁶ samples per problem, the largest 41B model produced at least one solution passing the example tests for over 90% of problems. Those survivors were clustered down to at most 10 submissions, which achieved an average ranking in the top 54.3% of human participants on Codeforces (Li et al., 2022).
Read that as an accounting statement rather than a headline.
The example-test filter rejected more than 99% of generated samples before submission. That is an enormous amount of search removed by a cheap deterministic check.
But passing the example tests did not establish correctness. Those tests were a filter, not the final oracle. AlphaCode still clustered and reranked the survivors, and the submitted programs were judged by Codeforces against tests the system had not used for filtering. The model supplied candidate diversity; the example tests supplied a powerful rejection signal.
Bound it: competitive programming, self-contained problems, example tests handed to you in the problem statement, and a compute budget no ordinary project will spend. A million samples is not a strategy anyone should copy. The ratio is the lesson, not the method.
And the lesson is practical: improve the check as deliberately as you improve the generator. A stronger model cannot repair a verifier that accepts the wrong property. Later chapters measure both sides of that statement: Part 5 tests whether extra model diversity buys verified coverage, while Chapter 28 preserves an accepted-but-wrong result produced by an inadequate deterministic checker.
Recent benchmarks make the same asymmetry explicit. BrowseComp is built from 1,266 questions whose answers are hard to find but short and easy to check; its question writers had to confirm that GPT-4o, with and without browsing, OpenAI o1 and an early deep-research model could not solve them (OpenAI, 2025). Trainers working on questions they had not written solved 29.2% within a two-hour limit. GPT-4o scored 0.6%, GPT-4o with browsing 1.9%, o1 9.9%, and the deep-research agent 51.5%. It is a vendor benchmark of short answers, not a workplace claim, but it states this chapter’s asymmetry cleanly: generation expensive, checking cheap — provided someone manufactured the check.
PaperBench shows what manufacturing the check costs: replicating 20 ICML 2024 papers is graded against rubrics broken into 8,316 individually gradable tasks, co-developed with each paper’s authors, and the automated judge is validated on a benchmark of its own; the best agent averaged 21.0% (Starace et al., 2025). SWE-Bench Pro pushes coding evaluation toward longer work: 1,865 human-verified problems from 41 repositories, some expected to take a professional hours or days (Deng et al., 2025).
The through-line is that each of these required enormous effort spent building the verifier, before any generator could be measured against it.
Reading terrain
Now the general question, which is what the rest of this book is for.
Anywhere a deduction has to be made over text or recorded information, there is an opening — writing, research, adjudication, review, fraud detection. That instinct is right, and it is the first thing to look for.
But text and deduction alone get you a demo. Text, deduction, and a cheap check get you a process. A process that keeps its verdicts gets you one that can improve. That difference is the entire subject of this book, and it is why the terrain survey has six questions rather than one.
| Survey question | Software’s answer | If the answer is no |
|---|---|---|
| Is the work already expressed in text or symbols? | Tickets, code, docs, reviews | The steamroller has no purchase. Your first job is codification, not model calls. |
| Does it decompose into independently attemptable units? | One ticket, one diff | You cannot localize a failure or retry cheaply. Everything is one big bet. |
| Is there a verifier cheaper than doing the work? | Compiler, types, tests, CI | You have a demo, not a process (Chapter 21). |
| Is a wrong proposal cheap to discard? | Branch it, revert it | Every attempt is a production incident. You need authority gates before you need intelligence (Chapter 20). |
| Are attempts and verdicts recorded, including failures? | Issue history, CI logs, reverts | Every improvement starts from memory. You cannot tell whether a change helped. |
| Is there volume, with variation? | High volume with variation | Low volume: not worth a process. No variation: write a script (Chapter 3). |
The verifier row is load-bearing. It is the one people skip, and skipping it is the single most reliable way to build something that demos beautifully and cannot be deployed. The record row is the one people get half right: they keep what shipped and throw away what did not.
Apply the survey honestly to the openings listed above and they do not all score the same:
- Research and literature synthesis. Text, yes. Decomposes per claim, yes. Some checks are cheap and mechanical: does the cited source exist, does the quoted span occur, does the bibliographic identity match? Whether the source actually supports the claim can still require judgment. This is useful terrain precisely because mechanical checks and judgment can be separated.
- Adjudication, claims handling, compliance review. Text, yes. Decomposes per case, often. But whether a rule applies to recorded facts may itself require interpretation, and appeal outcomes are delayed evidence rather than automatic ground truth. The domain may still be tractable, but it is authority-heavy and verifier quality has to be established rather than assumed.
- Evaluation and grading. Text, yes. A rubric can make criteria explicit, but inter-rater agreement is evidence about consistency, not proof that the grade is correct. Chapter 21 returns to the danger of replacing an oracle with another judgment.
- Fraud detection. Text, yes. Volume, enormous. But the verifier is delayed and adversarial: you may learn you were wrong months later, if ever, and the terrain changes in response to what you deploy. This is the case that looks attractive on a naive reading and becomes difficult when the verifier row is taken seriously.
Keep that last one. A survey that returns “yes” for everything is not a survey.
You are not holding a chatbot
Here is the turn, and it is the reason this book exists.
Almost everyone is looking at the chat window and asking whether the thing behind it is smart. Is it conscious, is it reasoning, did it really understand the question, will it pass some exam. That is an interesting argument and it is the wrong object.
The chat window is a demonstration interface and a useful conversational surface. The mistake is treating the transcript as the process.
Software’s advantage was not that its information was merely text-shaped. Its work was already surrounded by task identities, versioned artifacts, executable checks, and durable histories.
A chat transcript may be stored, but it does not by itself give you those records. It does not establish which unit of work an answer belongs to, what counted as done, which state was changed, which check passed, or which discarded attempts mattered. That is the first half of an answer to the question Chapter 8 asks — if the capability is this strong, where are the finished projects?
Build your application the way Scrum built software
Which produces the conclusion that took me longest to accept, and that most reorders what you do on Monday morning:
In a new domain, the first job is usually not to call a model. It is to build what Scrum built — and to build your application so that it keeps those books on itself from the first call.
Read the old ceremony one more time, this time as a design brief for software that uses AI:
| The practice | The rule for your application | Where this book builds it |
|---|---|---|
| Cut the work into tickets | Every piece of AI work is a unit with its own identity, not a message in a thread | Chapter 11 |
| Write down what done means before starting | Acceptance criteria are data the process checks, not a feeling at the end | Chapter 14 |
| Work in short cycles, each checked | One bounded proposal at a time, verified against the state it produced | Chapter 21 |
| Record what failed, not only what shipped | Failed calls, rejected claims and discarded attempts are kept with their reasons | Chapters 16–18 |
| Keep everything, good, bad and indifferent | Raw output is preserved before interpretation, so it can be judged again later | Chapter 17 |
| Retrospect, then change how you work | Change the process from the record, and check that the change helped | Chapters 27 and 30 |
Units. Descriptions of what is wanted. Acceptance criteria. A record of what was attempted and what happened, whichever way it went. A way to tell finished from unfinished. You produce the terrain before you roll it — and you keep producing it while you roll.
That is precisely what the rest of this book constructs: a machine for manufacturing the bookkeeping that makes a domain rollable, for domains that never had a Scrum.
The road consumes itself
One honest complication, and it is the most interesting finding in the chapter.
Del Rio-Chanona, Laurentsyeva, and Wachs measured what happened to Stack Overflow after ChatGPT’s release, using a difference-in-differences design against Russian and Chinese counterpart sites where ChatGPT access is limited, and against mathematics forums where the model was less capable. They estimate a 16% decrease in weekly posts on Stack Overflow, an effect that increases in magnitude over time and is larger for the most widely used programming languages. Posts made after ChatGPT receive similar voting scores to those before, so this is not merely the displacement of duplicate or low-quality content (del Rio-Chanona, Laurentsyeva & Wachs, 2023).
The pattern is suggestive: the estimated decline is larger for widely used programming languages, and the authors interpret the results as consistent with users substituting private model interactions for public question-and-answer activity where models are more useful.
Bound it: one platform, a specific window after release, a difference-in-differences estimate rather than a randomized experiment, and comparison platforms that differ from Stack Overflow in more ways than ChatGPT access. The study measures a decline in public posting; it does not measure the quality or future training value of the knowledge that was never posted.
The consequence for this book is therefore an inference, not another measured result. If more problem solving moves from public repositories into private interactions, the records your own process keeps become more valuable as organizational evidence. They do not replace the public commons, but they preserve what your organization tried, what was checked, what failed, and what later turned out to hold. That is a concrete reason to build the record-keeping of Parts 2 and 3.
Do this now
Twenty minutes, then two weeks. Run the terrain survey on your own domain, then start the record.
Take the work you actually want to automate — not a toy — and answer the six questions honestly in writing:
- Is it already expressed in text or symbols?
- Does it decompose into independently attemptable units?
- Is there a verifier cheaper than doing the work? Name it. If you cannot name it, write “none”.
- Is a wrong proposal cheap to discard?
- Are attempts and verdicts recorded, including failures?
- Is there volume, with variation?
A worked yield calculation grounds the survey, using only numbers already cited above: SWE-bench kept 2,294 instances from roughly 90,000 pull requests, so 2294 / 90000 ≈ 0.0255, about 2.5% of that initial pool became benchmark instances. Do not read 2.5% as a target or as the “cost of a strict verifier.” The construction applied several filters for suitability, installability, tests, and task behavior; the ratio describes this benchmark pipeline, not a general law.
Apply the same idea to your own shelf without inventing a pass threshold. Pull your last 40 work items and count how many have all three — a written request, a check cheaper than the work, and a recorded outcome:
| Count | Meaning |
|---|---|
| Complete triples / 40 | Your current harvest yield |
| Missing written request | Work that cannot yet be replayed from an explicit task |
| Missing cheap check | Work whose verification cost is still unresolved |
| Missing recorded outcome | Work that cannot yet become a labeled evaluation example |
Whether that is enough for a useful pilot depends on the task diversity, the verifier, and what you intend to measure; this chapter has not established a universal minimum.
Then start the record. For the next two weeks, every time you use AI on that work, write one line: what the attempt was for, what came back, what checked it, and the verdict — including attempts you discarded, with the reason. The goal is not to manufacture a balanced training set. It is to stop losing the evidence needed to tell what happened.
You can now do one new thing: price the verifier before pricing the model, and choose between designs — build bookkeeping first, or admit the result will be a demo.
If row 3 says “none”, stop and answer a different question first: what would a verifier for this look like, and what would it cost to build? That answer, not a prompt, is your first increment — and Chapter 7 turns it into something you can run.
Failure modes
- Choosing a domain by difficulty instead of checkability. Hard-but-checkable beats easy-but-uncheckable.
- Skipping the verifier row of the survey. Produces a system that demos perfectly and can never be trusted in production.
- Keeping only the successes. A record with no failures in it cannot tell good from bad, and cannot show whether a change helped.
- Adopting the ceremony without the record. Stand-ups and sprints with no written acceptance criteria and no verdicts produce nothing anyone can check against.
- Buying a better model to fix a weak verifier. A stronger generator does not make an inadequate check adequate. AlphaCode shows the leverage of aggressive filtering; Chapter 28 later shows that a deterministic check can still accept the wrong property.
- Treating another model’s agreement as the check. Covered properly in Chapter 21; noted here because agreement is not independent evidence merely because it came from another call.
- Assuming the public record stream is static. The Stack Overflow study measured a post-ChatGPT decline in public posting, especially in widely used languages. It did not measure the quality or future training value of the missing posts.
- Calling a model before the terrain exists. In an uncodified domain, the first increment is units, descriptions, and acceptance criteria — not a prompt.
What this chapter established
- Software became an unusually favorable early domain for AI because its work already had bounded tasks, versioned changes, cheap automated checks, and durable records. The claim is about that structure, not that Scrum caused language-model capability.
- The historical parallel is useful but bounded: machine learning benefits from objectives and feedback, while software engineering independently accumulated tasks, attempts, checks, and outcomes that later systems could reuse.
- The receipt is visible in SWE-bench’s construction: roughly 90,000 pull requests filtered into 2,294 task instances built from issues, patches, executable environments, and discriminating tests.
- Failures are useful evidence when they remain tied to the task and the check. SWE-Gym’s verifier-training experiment used 1,318 passing and 1,318 failing trajectories; an unchecked attempt does not establish whether the work succeeded.
- The operative property is checkability, not logicality. Compilers, type systems, tests, CI, and reversible versioned changes give software unusually cheap evidence about declared properties without proving general correctness.
- AlphaCode’s example-test filter removed over 99% of generated samples. That shows the leverage of cheap rejection, not that the example tests themselves supplied all correctness.
- Terrain survey: codified, decomposable, cheaply checkable, cheap to be wrong, recorded including failures, and high-volume-with-variation. The verifier row is the one that most often changes an attractive demo into a harder engineering problem.
- A chat transcript is not process state. Build explicit units, criteria, attempts, verdicts, and durable records around the model rather than asking the conversation itself to carry them.
- The Stack Overflow study measures a decline in public posting consistent with substitution toward private model interactions. The book’s inference is that preserving your own task, attempt, check, and outcome history becomes more important when useful work increasingly happens in private.
Next
We know which terrain is worth rolling, and we know the architecture: a confined stochastic proposal step behind a deterministic check, with every attempt and verdict kept. In most domains, though, the deterministic check bottoms out somewhere in a person — no compiler exists for “is this claim adequately supported.”
That person is the last verifier in the chain, and the last verifier degrades. Not through laziness, and not in a way that discipline fixes. The next chapter is about the specific way it happens, why success causes it rather than failure, and what has to be true of your process for review to survive contact with a system that is usually right.
Continue with Meat Proxy.
References
- David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning Representations by Back-Propagating Errors. Nature, vol. 323 (1986), pp. 533–536. https://doi.org/10.1038/323533a0
- Craig Larman and Victor R. Basili. Iterative and Incremental Development: A Brief History. Computer, vol. 36, no. 6 (2003), pp. 47–56. https://doi.org/10.1109/MC.2003.1204375
- Hirotaka Takeuchi and Ikujiro Nonaka. The New New Product Development Game. Harvard Business Review, January 1986, pp. 137–146. https://hbr.org/1986/01/the-new-new-product-development-game
- Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? International Conference on Learning Representations (ICLR), 2024. https://arxiv.org/abs/2310.06770
- Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang. Training Software Engineering Agents and Verifiers with SWE-Gym. International Conference on Machine Learning (ICML), 2025. https://arxiv.org/abs/2412.21139
- Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, et al. Competition-Level Code Generation with AlphaCode. Science, vol. 378, no. 6624 (2022), pp. 1092–1097. https://doi.org/10.1126/science.abq1158
- OpenAI. BrowseComp: A Benchmark for Browsing Agents. arXiv:2504.12516, April 2025. https://arxiv.org/abs/2504.12516
- Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. PaperBench: Evaluating AI’s Ability to Replicate AI Research. arXiv:2504.01848, 2025. https://arxiv.org/abs/2504.01848
- Xiang Deng et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? arXiv:2509.16941, 2025. https://arxiv.org/abs/2509.16941 — 1,865 problems, 41 repositories.
- Maria del Rio-Chanona, Nadzeya Laurentsyeva, and Johannes Wachs. Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow. arXiv:2307.07367, 2023. https://arxiv.org/abs/2307.07367