Debugging AI cover
AIBussin Book In development

Debugging AI

Progress from deterministic debugging to diagnosing models, evidence, trajectories, and AI-generated work.

Debugging AI builds one continuous debugging system from Python tracebacks to production AI incidents.

The progression is deliberate:

  1. Debug values and state in deterministic software.
  2. Cross the bridge: debug the interactive and numerical systems in between — notebooks, data, tensors, training, and evaluation — where state accumulates invisibly and every result must be reproduced before it is trusted.
  3. Debug distributions and uncertainty in learned systems.
  4. Debug intent and evidence in prompts, retrieval, and hallucinations.
  5. Debug trajectories in agents.
  6. Debug causes experimentally and convert diagnoses into prevention.

The debugging objects underneath that progression stay fixed — values, state, distributions, evidence, trajectories. Step 2 is not a new object; it is state stretched from a single binding to an entire accumulated computation, and the on-ramp to the distributional thinking step 3 depends on.

By the end, the reader should be able to move from a symptom to a scoped, reversible, causally-supported diagnosis, design discriminating experiments, and convert confirmed failures into durable regression artifacts.

The five objects are deep-structure categories, not a taxonomy of convenience: skilled diagnosis increasingly depends on classifying the structure of the failure — a value, an accumulated state, a distribution, an evidence chain, a trajectory — rather than reacting to its surface symptom, and each object selects its diagnostic method.

A diagnostic root-cause account, in this book, is an evidence-backed, reversible explanation sufficient to predict the observed failure within a defined diagnostic envelope. It is built by climbing a ladder — a difference, then a relevant difference, then a difference that changes the outcome when changed, then one that survives replication and reversal. It is not a claim that every system exposes one objective earliest cause: where ordering, observability, or intervention break down, the honest output is PROVISIONAL or INCONCLUSIVE, and socio-technical incidents carry a separate list of contributing factors rather than a single root.

Contents

Chapters

02

The First Divergence

Chapter 2 develops the first divergence within the book's end-to-end AI debugging system.

04

The Debugging Stack

Chapter 4 develops the debugging stack within the book's end-to-end AI debugging system.

07

Debug the Boundary

Chapter 7 develops debug the boundary within the book's end-to-end AI debugging system.

09

Environment Bugs

Chapter 9 develops environment bugs within the book's end-to-end AI debugging system.

11

Hidden Notebook State

Chapter 11 develops hidden notebook state within the book's end-to-end AI debugging system.

12

Reproducible Notebooks

Chapter 12 develops reproducible notebooks within the book's end-to-end AI debugging system.

15

When Training Goes Wrong

Chapter 15 develops when training goes wrong within the book's end-to-end AI debugging system.

16

Debugging Evaluation

Chapter 16 develops debugging evaluation within the book's end-to-end AI debugging system.

22

Internal Signals

Chapter 22 develops internal signals within the book's end-to-end AI debugging system.

25

Debugging Intent

Chapter 25 develops debugging intent within the book's end-to-end AI debugging system.

28

Debugging AI Research

Chapter 28 develops debugging ai research within the book's end-to-end AI debugging system.

29

Debugging Coding Agents

Chapter 29 develops debugging coding agents within the book's end-to-end AI debugging system.

31

Minimize the Prompt

Chapter 31 develops minimize the prompt within the book's end-to-end AI debugging system.

32

Retrieval Is a Pipeline

Chapter 32 develops retrieval is a pipeline within the book's end-to-end AI debugging system.

34

Debugging Hallucinations

Chapter 34 develops debugging hallucinations within the book's end-to-end AI debugging system.

36

An Agent Is a Trajectory

Chapter 36 develops an agent is a trajectory within the book's end-to-end AI debugging system.

37

Trace the Agent

Chapter 37 develops trace the agent within the book's end-to-end AI debugging system.

38

Agent Failure Taxonomy

Chapter 38 develops agent failure taxonomy within the book's end-to-end AI debugging system.

41

Causal Replay

Chapter 41 develops causal replay within the book's end-to-end AI debugging system.

42

Trajectory Diff

Chapter 42 develops trajectory diff within the book's end-to-end AI debugging system.

43

Multi-Agent Systems

Chapter 43 develops multi-agent systems within the book's end-to-end AI debugging system.

45

The AI Crash Dump

Chapter 45 develops the ai crash dump within the book's end-to-end AI debugging system.

46

Diagnostic AI Invariants

Chapter 46 develops diagnostic ai invariants within the book's end-to-end AI debugging system.

50

AIDebugBench

Chapter 50 develops aidebugbench within the book's end-to-end AI debugging system.

51

Debug the Debugger

Chapter 51 develops debug the debugger within the book's end-to-end AI debugging system.

52

AI Observability

Chapter 52 develops ai observability within the book's end-to-end AI debugging system.

56

Debugging in Production

Chapter 56 develops debugging in production within the book's end-to-end AI debugging system.

57

The Ten-Minute Debug

Chapter 57 develops the ten-minute debug within the book's end-to-end AI debugging system.

60

The Debugging AI Toolkit

Chapter 60 develops the debugging ai toolkit within the book's end-to-end AI debugging system.