Training: Which Link in the Learning Chain Is Broken?
Diagnose models that run but do not learn by walking TASK → OBJECTIVE → DEPENDENCY → GRADIENT → OPTIMIZER OWNERSHIP → PARAMETER UPDATE → CAPABILITY, using one fixed batch, exact parameter identity, controlled SGD, and tiny-batch overfitting to find the first missing consequence of learning.
Here is an ordinary piece of transfer-learning code. A small model, a frozen feature extractor, an optimizer, and a new classification head for the new task.
model = TinyClassifier()
for p in model.features.parameters():
p.requires_grad_(False)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-2, weight_decay=0.0)
model.head = nn.Linear(32, 2) # adaptation, performed later
Train it for two hundred steps on a balanced, perfectly aligned, entirely learnable binary task:
step= 0 loss=0.683606 acc=0.7144
step= 50 loss=0.683606 acc=0.7144
step= 100 loss=0.683606 acc=0.7144
step= 150 loss=0.683606 acc=0.7144
step= 200 loss=0.683606 acc=0.7144
Not approximately constant. Identical to six decimal places, two hundred times.
Nothing raised. Now look at the evidence a normal debugging session would collect:
loss 0.683606
loss.requires_grad True
loss.grad_fn NllLossBackward0
head.weight grad is None=False finite=True norm=1.555253e-01
head.bias grad is None=False finite=True norm=4.327927e-03
optimizer.step() returned without raising
The loss is finite. It has a grad_fn. backward() populated a finite, nonzero gradient on the head. optimizer.step() ran. Every check a programmer normally reaches for reports health, and the model is completely stationary:
features.0.weight delta=0.000000e+00
features.0.bias delta=0.000000e+00
features.2.weight delta=0.000000e+00
features.2.bias delta=0.000000e+00
head.weight delta=0.000000e+00
head.bias delta=0.000000e+00
One check tells you why, and it is not a check about gradients:
features.0.weight requires_grad=False in_optimizer=True
features.0.bias requires_grad=False in_optimizer=True
features.2.weight requires_grad=False in_optimizer=True
features.2.bias requires_grad=False in_optimizer=True
head.weight requires_grad=True in_optimizer=False
head.bias requires_grad=True in_optimizer=False
The optimizer was constructed one line too early. It holds the four frozen feature tensors, which have no gradient and are skipped, and the two parameter objects belonging to the old head, which no longer participates in the forward pass. The new head is registered on the model, listed by named_parameters(), saved by state_dict(), reached by autograd, and absent from the optimizer. So optimizer.step() faithfully updates nothing.
Two statements are doing all the damage here, and both of them are things people say without noticing:
loss.backward()proves that a derivative was computed. It does not prove that the value receiving that derivative will ever change.
optimizer.step()proves that an optimizer method ran. It does not prove that it owned the parameters you intended to train.
Where we are
Chapter 10 finished attention with a very specific and very limited certificate. It established that an implementation has correct external shapes, correct axis semantics, correct head rearrangement, correct mask semantics, the right softmax axis, verified causal behavior, agreement with PyTorch’s own reference implementation, and finite gradients reaching every expected parameter.
Then it stopped, deliberately, and handed this chapter the harder case:
forward runs
loss is finite
backward runs
gradients exist
optimizer.step() runs
and the model is still useless
The opening above is exactly that state. So is a model whose loss sits at log(num_classes) forever, a model whose accuracy never leaves the majority baseline, a model that trains beautifully on garbage, and a model that becomes NaN on step three.
Chapter 3 gave us the machinery for one part of this: which parameters have a path to the loss. Chapter 5 gave us another: which structure actually owns a tensor object. Chapter 7 gave us a third: whether the representation the model sees means what we intended. This chapter is where those become one investigation, because a training system is not one operation that either works or does not.
What is the first consequence of learning that fails to appear?
The environment
Every number, table and trace in this chapter came from executing the code shown, with fixed seeds, in this environment:
PyTorch 2.13.0
Python 3.12.3
OS Linux, CPU only
There is no CUDA device here, so this chapter makes no claims about mixed precision or device-specific numerics. Those belong to Chapter 12 and are left there.
The technique: prove the chain one boundary at a time
Every chapter has added an investigation method. Chapter 2 asked for the first wrong tensor rather than the first illegal one. Chapter 3 asked where the gradient path stops existing. Chapter 5 asked which structure disagreed about ownership. Chapter 7 asked where a sample stopped satisfying its contract. Chapter 8 asked where derived and observed geometry first diverged. Chapter 10 asked which attention invariant failed first.
Chapter 11 adds:
Trace one fixed batch from the task contract to parameter movement. Stop at the first expected consequence that fails to occur.
The reason this works is that learning is not an event. It is a chain of causes, and every arrow in it implies evidence that can be measured on one batch:
sample + target
↓ does y still describe x?
forward computation
↓ are the outputs in the form the objective expects?
objective / loss
↓ is this scalar the loss we intended?
autograd dependency
↓ does that loss depend on the parameters we mean to train?
gradient on intended parameters
↓ did those parameters receive finite derivatives?
optimizer ownership
↓ does the optimizer hold those exact objects?
parameter update
↓ did those exact objects move?
changed model behavior
↓ can repeated controlled steps fit a tiny fixed problem?
loss reduction on a controlled problem
The opening failure satisfies every one of those boundaries up to ownership, and fails there. A detached hidden layer gets as far as gradients. A shuffled target relation fails at the very first one and satisfies every other. Those three failures produce the same headline symptom — a loss that will not move usefully — and the same reflex, which is to change the learning rate.
The learning rate is not on that list. It appears later in this chapter, once every boundary above it has evidence, because tuning an optimizer against a broken task contract is the single most common way to waste a week.
The important shift is from:
“My model isn’t learning. What hyperparameter should I change?”
to:
“Which link in the learning chain have I actually proved?”
The controlled problem
Debugging without a baseline is guessing with extra steps, so this chapter builds one small system and keeps it for the whole chapter. Every deliberate failure below differs from it by exactly one change.
import torch
import torch.nn as nn
import torch.nn.functional as F
RULE_W = torch.tensor([1.0, 0.75])
def known_rule(x):
return ((x @ RULE_W) > 0).long()
def make_data(n=2048, seed=0):
g = torch.Generator().manual_seed(seed)
x = torch.randn(n, 2, generator=g)
return x, known_rule(x)
class TinyClassifier(nn.Module):
def __init__(self, hidden=32, num_classes=2):
super().__init__()
self.features = nn.Sequential(
nn.Linear(2, hidden), nn.ReLU(),
nn.Linear(hidden, hidden), nn.ReLU(),
)
self.head = nn.Linear(hidden, num_classes)
def forward(self, x):
return self.head(self.features(x))
def build(seed=42, **kw):
torch.manual_seed(seed)
return TinyClassifier(**kw)
The task is linearly separable by a rule we wrote down: class 1 when x0 + 0.75*x1 > 0. That single property is what makes the chapter possible. Because the rule is known, target alignment is checkable exactly rather than by eye. Because the problem is easy, a failure to learn is unambiguous rather than a matter of patience. And because the whole thing runs on CPU in under a second, every claim below could be re-derived by a reader in a terminal.
The model is deliberately split into features and head, because that is the structure in which the opening bug actually occurs in real code.
The healthy reference
train class counts: [1039, 1009]
majority baseline : 0.50732421875
chance loss ln(2) : 0.6931471824645996
val class counts : [258, 254] val majority 0.50390625
model = build(42)
opt = torch.optim.AdamW(model.parameters(), lr=1e-2, weight_decay=0.0)
for step in range(301):
opt.zero_grad(set_to_none=True)
loss = F.cross_entropy(model(x), y)
loss.backward()
opt.step()
step train loss train acc val loss val acc
0 0.6933 0.4976 0.6938 0.4922
50 0.0158 0.9990 0.0121 0.9961
100 0.0079 0.9995 0.0083 0.9961
150 0.0045 0.9995 0.0075 0.9961
200 0.0028 0.9995 0.0060 0.9961
250 0.0019 1.0000 0.0049 0.9980
300 0.0014 1.0000 0.0042 0.9980
Three numbers here are worth keeping, because later comparisons are meaningless without them. The initial loss is 0.6933, which is close to ln(2): the cross-entropy loss produced by equal probability on two classes. An arbitrary untrained classifier is not guaranteed to emit uniform probabilities; the near-match here tells us that this particular initialization begins close to that reference. The majority baseline is 0.5073, which is what “always predict the larger class” achieves on this dataset. And a healthy run reaches 0.0014 and perfect training accuracy in three hundred full-batch steps.
That last number matters most. When a broken variant plateaus at 0.68, the interesting fact is not that 0.68 is a large number. It is that this exact model, on this exact data, with this exact optimizer, reached 0.0014.
A note on the baseline, because it is routinely over-read. The majority baseline tells you what a metric looks like when the model has learned nothing useful. It does not tell you why the model is sitting there, and a model can beat it while learning nothing at all — the opening failure sat at 0.7144 accuracy, well above 0.5073, purely because a random projection followed by a random linear head happened to correlate with a linear rule. The baseline is a symptom threshold. It is not a diagnosis.
Boundary A: does the target still describe the input?
Chapter 7 established that preprocessing is part of the model’s definition, and that a representation can be perfectly legal and completely wrong. It also established the method: state the contract, then check it, independently of whether optimization succeeds. This chapter reuses that method rather than reteaching the five transform categories.
For a classification batch, the contract has a structural part and a semantic part.
def task_report(x, y, num_classes, rule=None):
print(f" x shape={tuple(x.shape)} dtype={x.dtype} "
f"finite={torch.isfinite(x).all().item()}")
print(f" range=[{x.min().item():+.3f}, {x.max().item():+.3f}] "
f"mean={x.mean().item():+.4f} std={x.std().item():.4f}")
print(f" y shape={tuple(y.shape)} dtype={y.dtype} "
f"min={y.min().item()} max={y.max().item()}")
counts = torch.bincount(y, minlength=num_classes)
print(f" class counts {counts.tolist()} majority baseline "
f"{(counts.max() / counts.sum()).item():.4f}")
print(f" aligned rows x.shape[0]==y.shape[0] {x.shape[0] == y.shape[0]}")
if rule is not None:
print(f" known-rule agreement {(rule(x) == y).float().mean().item():.4f}")
Now destroy the alignment in the way real code destroys it — a permutation applied to one array and not the other:
perm = torch.randperm(len(x), generator=g)
x_bad = x[perm]
y_bad = y # forgot to permute
HEALTHY
x shape=(2048, 2) dtype=torch.float32 finite=True
range=[-4.094, +4.101] mean=-0.0096 std=0.9934
y shape=(2048,) dtype=torch.int64 min=0 max=1
class counts [1039, 1009] majority baseline 0.5073
aligned rows x.shape[0]==y.shape[0] True
known-rule agreement 1.0000
SHUFFLED INPUTS, LABELS LEFT ALONE
x shape=(2048, 2) dtype=torch.float32 finite=True
range=[-4.094, +4.101] mean=-0.0096 std=0.9934
y shape=(2048,) dtype=torch.int64 min=0 max=1
class counts [1039, 1009] majority baseline 0.5073
aligned rows x.shape[0]==y.shape[0] True
known-rule agreement 0.5000
Every reported structural statistic is unchanged. The shape, dtype, range, class counts and row count cannot reveal a permutation because no value or example was removed; the displayed mean and standard deviation are unchanged as well. What changed is the pairing, and exactly one row can see it: agreement with the known rule fell to 0.5000, which is chance on this balanced two-class problem.
Now train on it and watch what a healthy mechanism does with a broken problem:
all trainable params owned: True
step-0 total grad norm 1.6257e-02 all finite True min delta 1.414e-02 max delta 2.985e-01
step= 0 loss=0.6936 train acc=0.4976
step= 100 loss=0.6625 train acc=0.5918
step= 200 loss=0.6501 train acc=0.6045
step= 300 loss=0.6410 train acc=0.6167
held-out acc on the real task: 0.568359375
The loss is falling. Training accuracy is rising. Gradients are finite, the optimizer owns everything, every parameter moves. This is a mechanically healthy training system fitting an arbitrary input-label pairing. After 300 steps it has reached only 0.6167 training accuracy, so the experiment does not yet establish that all 2,048 random pairings were memorized; it establishes something more important here: optimization progress can coexist with a broken task relation. Held-out accuracy on the actual task is 0.568, barely above the 0.504 baseline.
That is why the task contract comes first:
A healthy training mechanism can optimize a broken learning problem and can look as though it is making progress while doing it.
Real datasets rarely come with a known_rule. But most of them come with something you can check cheaply: a class distribution you can predict, a handful of examples you can look at side by side, an invariant the label must satisfy, a subset whose answer you know. Chapter 7’s argument applies unchanged — a probe you build on purpose is the only instrument that sees this category, because none of the tensor metadata can.
The mirror image of misalignment is leakage: a feature that contains the answer. It fails at the same boundary and is caught by the same discipline, but it usually produces suspiciously good results rather than a model that will not learn, so it is not this chapter’s subject.
Boundary B: is this scalar the objective you intended?
The loss is a function, and a function is code, and code can be wrong while running perfectly. The cheapest possible check is to reproduce it by hand for one example.
logits = model(x[:1])
target = y[:1]
loss = F.cross_entropy(logits, target)
manual = -F.log_softmax(logits, dim=-1)[0, target.item()]
logits [[0.091391 0.048356]]
target 0
F.cross_entropy 0.67186153
-log_softmax[t] 0.67186153
allclose True
Then the batch version, which also pins down the reduction:
batch F.cross_entropy 0.69213694
mean of per-example manual 0.69213694
reduction is 'mean' True
Twelve lines that convert “I am using cross entropy” into “I know what number this function produced and why.” When it disagrees, you have found a boundary. When it agrees, you have ruled out one whole class of confusion for the price of one forward pass.
Chapter 7 already established the nuance about targets, and it is worth carrying rather than flattening. F.cross_entropy accepts two different target representations, and both are legitimate:
index-target loss 0.69213694
prob-target loss 0.69213694
smoothed prob target 0.69224751
label_smoothing=0.1 0.69224751
Class-index targets and one-hot probability targets produce exactly the same number here, and hand-built 0.95/0.05 soft targets reproduce label_smoothing=0.1 exactly, which for two classes is what smoothing by 0.1 means. So “cross entropy requires integer class indices” is not true. What is universal for this function is the other side:
The model’s contribution to cross entropy must be logits. The target may be indices or per-class probabilities.
The anchor failure: softmax before cross entropy
probs = logits.softmax(dim=-1)
loss = F.cross_entropy(probs, y)
This is the mistake everyone has made once. It runs. It produces a finite scalar. It has a grad_fn. Gradients flow, the optimizer owns everything, parameters move. Run it from an identical initialization against the correct version:
logits (correct)
step loss grad norm train acc
0 0.6933 3.0339e-01 0.4976
1 0.6467 3.0642e-01 0.9395
50 0.0158 1.2085e-02 0.9990
100 0.0079 5.2352e-03 0.9995
200 0.0028 1.8874e-03 0.9995
300 0.0014 4.1497e-03 1.0000
pre-softmaxed
step loss grad norm train acc
0 0.6932 1.5088e-01 0.4976
1 0.6695 1.5858e-01 0.9390
50 0.3220 5.8824e-03 0.9995
100 0.3178 3.1225e-03 0.9995
200 0.3149 1.8243e-03 0.9995
300 0.3141 1.0929e-03 1.0000
Read that carefully, because the naive story is wrong. The broken version learns. It reaches 100% training accuracy. What it does not do is reduce its loss below 0.3141, and it never will, because that is the floor of the objective it is actually optimizing:
perfect probs [1,0], target 0 -> 0.31326165795326233
-log(e^1/(e^1+e^0)) = 0.31326165795326233
Cross entropy applies a log-softmax to whatever it is given. Given probabilities in [0, 1], the best achievable input is a one-hot vector, whose log-softmax at the correct class is -log(e / (e + 1)), or about 0.3133. The loss can approach that and stop. A programmer watching this run sees a loss that plateaus at a suspiciously specific number and starts changing learning rates, when the correct observation is that the objective has a floor it is already at.
A finite, differentiable, decreasing loss is not evidence that it is the loss you intended.
Note also what this does to the gradient: the pre-softmax run has roughly half the gradient norm at step 0, because the softmax squashes the model’s output range before the loss ever sees it. Two objectives, same model, same initialization, same data, different gradients, and no error anywhere.
Boundary C: which parameters does the loss actually depend on?
The usual check is:
print("loss.requires_grad:", loss.requires_grad)
print("loss.grad_fn:", loss.grad_fn)
Keep it, and be precise about what it establishes. loss.requires_grad becomes True as soon as any single tensor on the path requires gradients. Chapter 3 made this concrete: a loss with a healthy grad_fn is evidence that something is connected, and no evidence at all about what.
Change one line in the model:
class DetachedClassifier(TinyClassifier):
def forward(self, x):
h = self.features(x)
h = h.detach() # the one changed line
return self.head(h)
loss 0.693254
loss.requires_grad True
loss.grad_fn NllLossBackward0
parameter req_grad grad in_opt delta
features.0.weight True None True 0.0000e+00
features.0.bias True None True 0.0000e+00
features.2.weight True None True 0.0000e+00
features.2.bias True None True 0.0000e+00
head.weight True 1.5752e-01 True 7.8739e-02
head.bias True 5.3857e-03 True 1.4142e-02
Both graph checks pass. Four of six parameters have requires_grad=True, are in the optimizer, and receive no gradient at all. This is Chapter 3’s question, asked at model scale:
Which intended parameter first stops having a path to this loss?
And here is what makes it dangerous rather than merely broken:
step= 0 loss=0.6845 acc=0.5649
step= 100 loss=0.3047 acc=0.9438
step= 200 loss=0.2122 acc=0.9590
step= 300 loss=0.1676 acc=0.9702
The model still reaches 97%. A linear head on a fixed random projection is enough for this task, so the failure does not announce itself as a failure. It announces itself as a model that is slightly worse than it should be — 0.1676 against the reference run’s 0.0014 — which is the kind of gap that gets attributed to hyperparameters and never investigated. The feature extractor is not dead weight: it still participates in every forward pass and supplies the fixed representation the head uses. What is dead is its learning path. Four trainable parameter tensors are carried through the model but can never adapt, and the only visible symptom may be a merely disappointing number.
This is a partial graph break, not “autograd is broken”, and the distinction is the whole diagnosis. Everything downstream of the detach is intact, which is exactly why editing anything downstream cannot help.
None is not zero, and zero is not None
Those two states look similar in a print statement and mean completely different things.
Here is a controlled way to produce the second. Drive one layer’s pre-activations far negative so its ReLU output is identically zero:
model = build(42)
with torch.no_grad():
model.features[0].bias.fill_(-100.0)
features.0 preactivation max : -96.8088
post-ReLU zero fraction : 1.0
parameter grad norm max|g|
features.0.weight tensor 0.0000e+00 0.0000e+00
features.0.bias tensor 0.0000e+00 0.0000e+00
features.2.weight tensor 0.0000e+00 0.0000e+00
features.2.bias tensor 1.4861e-02 7.9112e-03
head.weight tensor 1.5895e-02 4.5727e-03
head.bias tensor 3.6695e-02 2.5947e-02
Every one of those six parameters has a gradient tensor. Three of those tensors are exactly zero. The mechanism is local and readable: the ReLU derivative is zero wherever its input was negative, so nothing flows back into features.0; and features.2.weight’s derivative is proportional to its input, which is all zeros, so it is zero too. features.2.bias is unaffected, because a bias derivative does not depend on the layer’s input. The model can still change — through biases only — which is why this produces a slow, confusing, partial kind of learning rather than a clean stop.
Side by side with the detached model:
detached model features.0.weight.grad is None: True
dead-relu model features.0.weight.grad is None: False
Same headline symptom, different boundary, different repair.
Two things to avoid concluding from this. First, a zero gradient does not mean “dead ReLU” — that is one mechanism among several. Second, and more important, do not attach absolute numerical thresholds to gradient magnitudes. There is no architecture-independent “healthy gradient norm”, because the number depends on the loss scale, the parameter shapes, the depth, the initialization and the batch size. What is diagnostic is comparison: how a norm changes over steps, how it differs across layers, how it compares against a known-good run, and where non-finite values first appear.
Boundary D: gradient evidence
The useful gradient report is not a wall of statistics. It distinguishes four states per parameter, and where each state occurs in the model is the diagnosis:
missing grad is None
zero grad tensor exists, every element is zero
finite/nonzero grad tensor exists and carries signal
non-finite grad tensor contains NaN or Inf
Total gradient norm is worth computing, as one row rather than a section:
total = torch.stack([p.grad.norm() for p in model.parameters()
if p.grad is not None]).norm()
On the healthy reference at step 0 that is 3.4720e-01. That number means nothing on its own. It becomes evidence when the same measurement on the same model with a pre-softmaxed loss reads 1.7188e-01, or when it goes from 3.47e-01 to 5.94e+05 in one step, or when a layer near the input reads 0.0 while a layer near the output reads 1.6e-02.
One rule about the instrument itself, which the rest of the chapter depends on:
Observation and intervention are different operations. A diagnostic pass must not clip gradients, change the model’s mode, rebuild the optimizer, step a scheduler or alter regularization.
clip_grad_norm_ in particular does not belong in a debugging step, because it modifies the thing being measured. When it is used deliberately, know exactly what it returns:
signature: (parameters, max_norm, norm_type=2.0, error_if_nonfinite=False, foreach=None)
total grad norm before clipping : 0.347198
value returned by the function : 0.347198
total grad norm after clipping : 0.100000
returned value is the PRE-clip norm: True
The returned value is the norm before clipping. Logging it is genuinely useful; logging it while believing it describes the gradients the optimizer then consumed is not. And “clipping made the NaN go away” is a workaround, not a diagnosis.
Boundary E: does the optimizer own those exact objects?
Chapter 5 established the mechanism: an optimizer is a list of tensor objects captured at construction time, and membership is by object identity — not by name, not by shape, not by value. This chapter turns that into a two-way audit, because there are two ways for the two structures to disagree and only one of them is usually checked.
def optimizer_param_ids(optimizer):
return {id(p) for g in optimizer.param_groups for p in g["params"]}
def ownership_audit(model, optimizer):
"""Two-way membership by object identity, in both directions."""
opt_ids = optimizer_param_ids(optimizer)
model_ids = {id(p) for p in model.parameters()}
missing = [n for n, p in model.named_parameters()
if p.requires_grad and id(p) not in opt_ids]
stale = [p for g in optimizer.param_groups for p in g["params"]
if id(p) not in model_ids]
return missing, stale
missing is the fatal direction: trainable parameters the optimizer will never touch. stale is the diagnostic direction: tensors the optimizer still holds that the model no longer reaches. On the opening failure, missing is [head.weight, head.bias] and stale has two entries. Together, those facts strongly suggest model surgery after optimizer construction: the model has new trainable objects while the optimizer still owns unreachable old ones. The object identities establish the mismatch; the matching head-like shapes are supporting context, not a unique signature.
The subtlety worth executing
The optimizer’s zero_grad clears gradients for the parameters the optimizer owns. The new head is not one of them, so nothing in the training loop ever clears its gradient:
iteration head.weight.grad norm features.0.weight.grad
1 0.155525 None
2 0.311051 None
3 0.466576 None
4 0.622101 None
5 0.777627 None
ratio to first iteration: 5.000010504061109
Exactly linear accumulation. The gradient is identical every iteration, because the parameters never move, so five backward passes deposit exactly five times the first gradient into a buffer nothing will ever read. And:
optimizer.zero_grad(set_to_none=True) then inspect:
head.weight.grad is None : False
Chapter 1 established that PyTorch accumulates into .grad and that something has to clear it. Chapter 5 established that optimizer.zero_grad() is that something, and that it can only clear what it was given. Here both facts combine into a parameter that silently grows a gradient buffer for the entire run and never uses it.
For reference, zero_grad on this version defaults to set_to_none=True:
(self, set_to_none: bool = True) -> None
after default zero_grad, grad is None: True
after set_to_none=False, grad: tensor([[0., 0.], [0., 0.]])
That default matters for the previous section: after a default zero_grad, an owned parameter that receives no gradient has .grad is None, not a zero tensor. So “None” can mean “was cleared and never refilled” as well as “has no path to the loss”, and telling those apart is a question about when you looked.
The repair
Change one thing — construct the optimizer after the surgery:
model.head = nn.Linear(32, 2)
optimizer = torch.optim.AdamW(
[p for p in model.parameters() if p.requires_grad],
lr=1e-2, weight_decay=0.0,
)
features.0.weight requires_grad=False in_optimizer=False
features.0.bias requires_grad=False in_optimizer=False
features.2.weight requires_grad=False in_optimizer=False
features.2.bias requires_grad=False in_optimizer=False
head.weight requires_grad=True in_optimizer=True
head.bias requires_grad=True in_optimizer=True
features.0.weight delta=0.000000e+00
features.0.bias delta=0.000000e+00
features.2.weight delta=0.000000e+00
features.2.bias delta=0.000000e+00
head.weight delta=7.873873e-02
head.bias delta=1.414209e-02
step= 0 loss=0.675015 acc=0.7344
step= 50 loss=0.406795 acc=0.9297
step= 100 loss=0.302602 acc=0.9463
step= 150 loss=0.246815 acc=0.9570
step= 200 loss=0.210920 acc=0.9609
val acc 0.97265625
The frozen features still have delta=0, which is correct — that was the declared intent. The head moves, the loss falls, and held-out accuracy reaches 0.973. That is the difference between a stationary model and a working linear probe.
Two things changed in that construction and only one of them was necessary. Moving the optimizer after the surgery is the repair. Filtering on requires_grad is a separate, deliberate tightening: handing the frozen tensors to the optimizer would also have worked because parameters with .grad is None are skipped. But then optimizer membership would no longer mean “this is a parameter I intend to train.” Keeping the optimizer’s contents aligned with the declared trainable set makes later ownership reports easier to interpret.
Boundary F: did those objects actually move?
A gradient is a claim about a derivative. A parameter delta is a claim about the model. They are different measurements and the second one is the one that matters:
before = {n: p.detach().clone() for n, p in model.named_parameters()}
optimizer.step()
deltas = {n: (p.detach() - before[n]).norm().item()
for n, p in model.named_parameters()}
A parameter can have a perfectly healthy gradient and a delta of exactly zero if it is not owned by the optimizer, if its group’s learning rate is zero, or if the optimizer skipped it for its own reasons. The delta is the evidence.
But movement is not automatically the update you think it is. Stateful optimizers make the relationship between the current gradient and the current step indirect: momentum carries earlier gradients forward, Adam and AdamW maintain running moment estimates, and AdamW’s decoupled weight decay contributes a term that does not come from the gradient at all. So two questions have to stay separate:
DID THE PARAMETER MOVE?
measurable under the real optimizer, on any step
DID THIS GRADIENT PRODUCE EXACTLY THE UPDATE I EXPECT?
needs a controlled optimizer
Chapter 1, at model scale
Chapter 1 built training out of one line:
w_new = w - lr * grad
That was one scalar parameter and a hand-written loop. It is worth proving that the same equation still describes an actual torch.optim step on an actual neural network, because once you have seen it hold you stop treating the optimizer as an oracle.
The experiment has to be controlled, so it uses a copy of the model and a fresh plain SGD with every update modifier disabled:
probe = copy.deepcopy(model)
lr = 0.1
sgd = torch.optim.SGD(probe.parameters(), lr=lr,
momentum=0.0, weight_decay=0.0)
for p in probe.parameters():
p.grad = None
F.cross_entropy(probe(xb), yb).backward()
before = {n: p.detach().clone() for n, p in probe.named_parameters()}
grads = {n: p.grad.detach().clone() for n, p in probe.named_parameters()}
sgd.step()
parameter max|observed - (-lr*g)| ||delta|| lr*||g||
features.0.weight 2.724e-08 7.7473e-03 7.7473e-03
features.0.bias 2.725e-08 2.6625e-03 2.6625e-03
features.2.weight 7.422e-09 2.7025e-02 2.7025e-02
features.2.bias 6.636e-09 5.4703e-03 5.4703e-03
head.weight 7.218e-09 1.7827e-02 1.7827e-02
head.bias 2.794e-09 7.7633e-03 7.7633e-03
allclose over every parameter: True
w_new = w - lr * grad, elementwise, across 1,218 parameters in six tensors, to floating-point tolerance. The scalar loop from Chapter 1 was not merely toy arithmetic: under these deliberately restricted SGD settings, it is exactly what this optimizer does.
Now the same measurement under AdamW, from the same state, on its first step:
features.0.weight ||delta||=7.9999e-01 lr*||g||=7.7473e-03 ratio=103.260
features.0.bias ||delta||=5.6562e-01 lr*||g||=2.6625e-03 ratio=212.441
features.2.weight ||delta||=2.9306e+00 lr*||g||=2.7025e-02 ratio=108.439
features.2.bias ||delta||=5.5677e-01 lr*||g||=5.4703e-03 ratio=101.782
head.weight ||delta||=7.8740e-01 lr*||g||=1.7827e-02 ratio=44.170
head.bias ||delta||=1.4142e-01 lr*||g||=7.7633e-03 ratio=18.217
Between 18 and 212 times lr * ||grad||, and a different ratio for every tensor. That is not a bug; it is what an adaptive optimizer is for, and it is precisely why the equality above has to be stated with its conditions attached:
Δθ = -lr · gradis a property of the controlled plain-SGD experiment, not a universal optimizer invariant. Do not try to force AdamW into it — reproduce one step with fresh SGD instead.
This is also worth remembering when reading Chapter 7’s opening again: the same representation bug was catastrophic under SGD and nearly invisible under Adam. Adaptive update normalization changes what a gradient does to a parameter, which changes what a bug looks like from the outside.
One instrument: the learning ledger
Chapter 7 has a stage report. Chapter 8 has a shape ledger. Chapter 9 has a feature-space inspector. Chapter 10 has an attention ledger. This chapter’s instrument walks the learning chain over one fixed batch and names the first boundary whose evidence contradicts the intended training system.
The design constraints are the ones argued for above, with one important warning: this is not a read-only report. It calls backward() and optimizer.step(), so it mutates gradients, parameter values and optimizer state. Run it on a disposable copy/checkpoint, or treat the step as a deliberate sacrificial probe that you intend to keep. Within that probe it performs no clipping, scheduler step, mode change or regularization edit. It clears .grad on every model parameter before backward() so the gradient rows describe this batch rather than an earlier accumulation, but it reports any pre-existing gradients first so accidental carry-over remains visible.
The default expected set below is every parameter with requires_grad=True. That is correct for this dense controlled model. In conditional, sparse or mixture-style models, a trainable parameter may legitimately not participate in a particular batch; in that case the expected set must come from the model’s execution contract rather than from requires_grad alone.
def learning_ledger(model, optimizer, x, y, num_classes,
rule=None, loss_fn=F.cross_entropy):
"""One controlled, state-mutating probe over one fixed batch.
Calls backward() and optimizer.step(), so it changes gradients, model
parameters and optimizer state. It performs no clipping, scheduler step,
mode change or regularization edit. Gradients are cleared to None before
backward so the reported derivatives belong to this batch.
"""
rows, first = [], None
def row(section, label, text, ok=None, detail=None):
nonlocal first
rows.append([section, label, text, ok, detail])
if ok is False and first is None:
first = f"{section}/{label}"
named = dict(model.named_parameters())
expected = [n for n, p in named.items() if p.requires_grad]
# TASK ------------------------------------------------------------
row("TASK", "rows", f"x {x.shape[0]} / y {y.shape[0]}",
x.shape[0] == y.shape[0])
row("TASK", "types", f"x {x.dtype} / y {y.dtype}", y.dtype == torch.long)
row("TASK", "finite", f"x finite {torch.isfinite(x).all().item()}",
bool(torch.isfinite(x).all()))
counts = torch.bincount(y, minlength=num_classes)
row("TASK", "classes",
f"counts {counts.tolist()} majority baseline "
f"{(counts.max() / counts.sum()).item():.3f}",
bool(y.min() >= 0 and y.max() < num_classes))
if rule is not None:
agreement = (rule(x) == y).float().mean().item()
row("TASK", "known rule", f"agreement {agreement:.3f}", agreement > 0.99)
# FORWARD ---------------------------------------------------------
row("FORWARD", "mode", f"model.training = {model.training}")
logits = model(x)
finite = bool(torch.isfinite(logits).all())
row("FORWARD", "logits",
f"{tuple(logits.shape)} {logits.dtype} finite {finite}",
finite and logits.shape[0] == x.shape[0])
out = logits.detach()
row("FORWARD", "scale", f"mean {out.mean():+.4f} std {out.std():.4f} "
f"max|.| {out.abs().max():.4f}")
# LOSS ------------------------------------------------------------
loss = loss_fn(logits, y)
row("LOSS", "value", f"{loss.item():.6f}", bool(torch.isfinite(loss)))
if loss_fn is F.cross_entropy and y.dtype == torch.long:
manual = -F.log_softmax(logits[:1], dim=-1)[0, y[0]]
one = F.cross_entropy(logits[:1], y[:1])
row("LOSS", "one example",
f"manual {manual.item():.6f} vs reported {one.item():.6f}",
bool(torch.allclose(manual, one, atol=1e-6)))
row("LOSS", "graph",
f"requires_grad {loss.requires_grad} grad_fn "
f"{type(loss.grad_fn).__name__ if loss.grad_fn else None}",
loss.requires_grad and loss.grad_fn is not None)
# GRADIENT --------------------------------------------------------
carried = [n for n in expected if named[n].grad is not None]
row("GRADIENT", "pre-existing",
f"{len(carried)}/{len(expected)} already held a gradient before backward",
not carried)
for p in model.parameters():
p.grad = None
loss.backward()
absent = [n for n in expected if named[n].grad is None]
live = [n for n in expected if named[n].grad is not None]
nonfinite = [n for n in live if not torch.isfinite(named[n].grad).all()]
zero = [n for n in live if not (named[n].grad != 0).any()]
row("GRADIENT", "expected", f"{len(expected)} trainable parameters")
row("GRADIENT", "present", f"{len(live)}/{len(expected)}", not absent,
None if not absent else "None: " + ", ".join(absent))
row("GRADIENT", "finite", f"{len(live) - len(nonfinite)}/{len(live)}",
not nonfinite,
None if not nonfinite else "non-finite: " + ", ".join(nonfinite))
row("GRADIENT", "nonzero", f"{len(live) - len(zero)}/{len(live)}", None,
None if not zero else "exactly zero: " + ", ".join(zero))
if live:
total = torch.stack([named[n].grad.norm() for n in live]).norm().item()
row("GRADIENT", "total norm", f"{total:.4e}")
# OPTIMIZER -------------------------------------------------------
missing, stale = ownership_audit(model, optimizer)
groups = " ".join(
f"[{i}] lr {g['lr']:g} wd {g.get('weight_decay', 0):g} n {len(g['params'])}"
for i, g in enumerate(optimizer.param_groups))
row("OPTIMIZER", "groups", f"{type(optimizer).__name__} {groups}")
row("OPTIMIZER", "owned",
f"{len(expected) - len(missing)}/{len(expected)} trainable model parameters",
not missing, None if not missing else "missing: " + ", ".join(missing))
row("OPTIMIZER", "stale",
f"{len(stale)} optimizer tensors no longer reachable from the model",
not stale)
# UPDATE ----------------------------------------------------------
before = {n: p.detach().clone() for n, p in named.items()}
optimizer.step()
deltas = {n: (named[n].detach() - before[n]).norm().item() for n in expected}
still = [n for n, d in deltas.items() if d == 0.0]
row("UPDATE", "moved",
f"{len(deltas) - len(still)}/{len(deltas)} expected parameters",
not still, None if not still else "unchanged: " + ", ".join(still))
if deltas:
row("UPDATE", "delta range",
f"{min(deltas.values()):.3e} .. {max(deltas.values()):.3e}")
width = max(len(t) for _, _, t, _, _ in rows)
print(f"LEARNING LEDGER {type(model).__name__} {type(optimizer).__name__}"
f" batch {tuple(x.shape)}")
previous = None
for section, label, text, ok, detail in rows:
head = "" if section == previous else section
previous = section
mark = ""
if ok is False:
mark = (" <-- FIRST DIVERGENCE" if f"{section}/{label}" == first
else " fail")
print(f" {head:10s}{label:14s}{text:{width}s}{mark}")
if detail:
print(f" {'':10s}{'':14s}{detail}")
print(f" first divergence: {first or 'none'}")
return {"first_divergence": first, "loss": loss.item(), "deltas": deltas}
On the healthy reference:
LEARNING LEDGER TinyClassifier AdamW batch (256, 2)
TASK rows x 256 / y 256
types x torch.float32 / y torch.int64
finite x finite True
classes counts [115, 141] majority baseline 0.551
known rule agreement 1.000
FORWARD mode model.training = True
logits (256, 2) torch.float32 finite True
scale mean +0.1037 std 0.0457 max|.| 0.3531
LOSS value 0.694159
one example manual 0.671861 vs reported 0.671861
graph requires_grad True grad_fn NllLossBackward0
GRADIENT pre-existing 0/6 already held a gradient before backward
expected 6 trainable parameters
present 6/6
finite 6/6
nonzero 6/6
total norm 3.4720e-01
OPTIMIZER groups AdamW [0] lr 0.01 wd 0 n 6
owned 6/6 trainable model parameters
stale 0 optimizer tensors no longer reachable from the model
UPDATE moved 6/6 expected parameters
delta range 1.414e-02 .. 2.931e-01
first divergence: none
Twenty-two rows, one batch, and a claim about every boundary from the task contract to parameter movement. That is what “the training step is mechanically sound” should mean before anyone touches a hyperparameter.
Four broken systems, one instrument
Each of the systems below differs from the healthy reference by exactly one change, and each is run through the same function on the same batch. The later listings are abridged to the rows that differ; every omitted row is identical to the healthy run above.
Stale optimizer membership
TASK rows x 256 / y 256
types x torch.float32 / y torch.int64
finite x finite True
classes counts [115, 141] majority baseline 0.551
known rule agreement 1.000
FORWARD mode model.training = True
logits (256, 2) torch.float32 finite True
scale mean -0.0539 std 0.0429 max|.| 0.1612
LOSS value 0.685485
one example manual 0.675722 vs reported 0.675722
graph requires_grad True grad_fn NllLossBackward0
GRADIENT pre-existing 0/2 already held a gradient before backward
expected 2 trainable parameters
present 2/2
finite 2/2
nonzero 2/2
total norm 1.9480e-01
OPTIMIZER groups AdamW [0] lr 0.01 wd 0 n 6
owned 0/2 trainable model parameters <-- FIRST DIVERGENCE
missing: head.weight, head.bias
stale 2 optimizer tensors no longer reachable from the model fail
UPDATE moved 0/2 expected parameters fail
unchanged: head.weight, head.bias
delta range 0.000e+00 .. 0.000e+00
first divergence: OPTIMIZER/owned
Everything above OPTIMIZER reports health, including the gradient section, which is where most people stop looking. Two rows are worth reading closely. expected 2 trainable parameters is the ledger taking the freeze at its word: requires_grad=False is a declared intent, so frozen parameters are not counted as expected to move. And groups ... n 6 against owned 0/2 is the whole bug in two numbers — the optimizer holds six tensors and none of them is one of the two the model wants trained.
The UPDATE row is real but downstream. It is a consequence of the ownership failure, not an independent finding, which is exactly the discipline Chapters 8 and 10 established: fix the first divergence, then rerun.
Detached hidden representation
Every TASK, FORWARD and LOSS row is identical to the healthy run, so the listing starts where it stops being identical:
GRADIENT pre-existing 0/6 already held a gradient before backward
expected 6 trainable parameters
present 2/6 <-- FIRST DIVERGENCE
None: features.0.weight, features.0.bias,
features.2.weight, features.2.bias
finite 2/2
nonzero 2/2
total norm 1.9444e-01
OPTIMIZER groups AdamW [0] lr 0.01 wd 0 n 6
owned 6/6 trainable model parameters
stale 0 optimizer tensors no longer reachable from the model
UPDATE moved 2/6 expected parameters fail
unchanged: features.0.weight, features.0.bias,
features.2.weight, features.2.bias
first divergence: GRADIENT/present
The optimizer section is clean. Ownership is perfect. The failure is one boundary earlier, and the UPDATE row shows the same four names for a completely different reason: not “the optimizer does not hold them” but “they had no gradient to step with.”
Two failures, nearly identical UPDATE rows, opposite repairs. That is the argument for walking the chain in order rather than starting from the symptom.
Destroyed target alignment
TASK rows x 256 / y 256
types x torch.float32 / y torch.int64
finite x finite True
classes counts [115, 141] majority baseline 0.551
known rule agreement 0.434 <-- FIRST DIVERGENCE
FORWARD mode model.training = True
logits (256, 2) torch.float32 finite True
LOSS value 0.695320
one example manual 0.685739 vs reported 0.685739
graph requires_grad True grad_fn NllLossBackward0
GRADIENT present 6/6
finite 6/6
nonzero 6/6
total norm 1.1394e-01
OPTIMIZER owned 6/6 trainable model parameters
stale 0 optimizer tensors no longer reachable from the model
UPDATE moved 6/6 expected parameters
delta range 1.414e-02 .. 2.960e-01
first divergence: TASK/known rule
Every other row in the ledger reports health, under one broken row. If the ledger started at FORWARD — which is where a debugging session usually starts, because the model is the interesting part — it would report a flawless training system optimizing nonsense.
The failure the ledger cannot see
Run it with the pre-softmaxed objective:
LOSS value 0.693597
graph requires_grad True grad_fn NllLossBackward0
GRADIENT present 6/6
finite 6/6
nonzero 6/6
total norm 1.7188e-01
OPTIMIZER owned 6/6 trainable model parameters
UPDATE moved 6/6 expected parameters
first divergence: none
first divergence: none, on a model that will plateau at 0.3141 forever. This is not a defect to patch; it is the instrument being honest about its reach. Notice what it did not do: the one example row is absent, because the loss function is not F.cross_entropy and the ledger declines to verify a formula it does not know. It reports what it can measure and makes no claim about what it cannot.
Chapter 10 made the same point about the attention ledger reaching Levels 1 through 3 and not Level 4. Here the boundary is:
The ledger proves that the machinery is intact. Only you can say whether the objective it is optimizing is the one you meant.
The failure matrix
Every row below is an executed experiment, not an expectation:
| Failure | Task | Loss | Gradient | Optimizer | Delta | Tiny batch |
|---|---|---|---|---|---|---|
| stale replaced head | pass | pass | pass, 2/2 finite | fail, 0/2 owned | fail, 0/2 moved | flat at 0.68758 |
| detached hidden state | pass | pass | fail, 2/6 present | pass | fail, 2/6 moved | 0.149 vs 0.002 |
| shuffled target relation | fail, 0.434 | pass | pass | pass | pass | memorizes |
| softmax before CE | pass | floor at 0.3133 | pass | pass | pass | stops at 0.3136 |
| lr = 1e4 | pass | pass at step 0 | pass at step 0 | pass | moved, wrong scale | diverges to NaN |
Four of those five have a first divergence the ledger names on its own. The learning rate is the exception and is worth being precise about: on step 0 every ledger row passes, including UPDATE/moved, because the parameters do move. What is wrong is the size of the movement relative to the parameters, which the delta range row records but does not judge. That failure needs the escalation instruments below, and it is the reason the ledger reports evidence rather than verdicts wherever a threshold would have to be invented.
Boundary G: can this system fit a tiny fixed batch?
Once every boundary above is clean for a single step, one question remains, and it is the only one that tests them together:
Can this model, this objective and this update path reduce this loss on these exact examples?
The test is not “train and see”. It is a controlled environment in which learning should be trivially easy, so that failure is informative:
def tiny_batch_capability(model, x_small, y_small, optimizer=None,
steps=500, lr=0.5):
"""Controlled capability test on a fixed batch.
optimizer=None copies the model and builds fresh plain SGD over its
trainable parameters, answering 'can this MODEL fit these examples under
this simplified optimizer?'. Passing an optimizer uses and mutates the
supplied model/optimizer, answering 'can this TRAINING SYSTEM fit these
examples?'. The caller is responsible for any model-level regularization
such as dropout and for any augmentation applied before this function.
"""
was_training = model.training
if optimizer is None:
model = copy.deepcopy(model)
optimizer = torch.optim.SGD(
[p for p in model.parameters() if p.requires_grad],
lr=lr, momentum=0.0, weight_decay=0.0)
model.train()
initial = None
for step in range(steps + 1):
logits = model(x_small)
loss = F.cross_entropy(logits, y_small)
if step == 0:
initial = loss.item()
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
logits = model(x_small)
final = F.cross_entropy(logits, y_small).item()
acc = (logits.argmax(-1) == y_small).float().mean().item()
model.train(was_training)
return initial, final, acc
That optimizer argument is not a convenience. It is the point:
system optimizer initial final accuracy
healthy as built 0.69209 0.00227 1.0000
healthy fresh SGD 0.69209 0.00227 1.0000
frozen feats + new head as built 0.68758 0.68758 0.7188
frozen feats + new head fresh SGD 0.68758 0.14840 0.9688
detached representation as built 0.69209 0.14906 0.9688
healthy, RANDOM labels fresh SGD 0.69527 0.00080 1.0000
Look at the two rows for the stale-head system. With the real optimizer, the loss does not move at all: 0.68758 to 0.68758. With a fresh one, it drops to 0.148 and fits 31 of 32 examples. The model was always capable. The training system was not. A tiny-batch test that quietly rebuilds the optimizer has repaired the bug it was supposed to detect, and will report that everything is fine.
So the version of the test that people usually write — new model, new optimizer, small batch — answers a question about capacity. It is a real question, and it is not the question you have when a specific training run will not learn.
What the tiny-batch test proves
Passing it, with the real training system, supports all of:
the representation reaches the model
the objective can be reduced on these examples
the graph delivers useful gradients
the optimizer can change the model
the model has enough capacity to fit this batch under this setup
Now the last row of that table, which is why this must not become a ritual. Replace the labels with coin flips:
random labels agreement with known rule: 0.4375
step loss accuracy
0 0.69527 0.3750
25 0.51209 0.7500
100 0.40176 0.7812
500 0.11932 1.0000
2000 0.00080 1.0000
Perfect memorization of 32 examples whose labels mean nothing. The tiny-batch test passed with flying colors on data that has no learnable structure whatsoever.
Passing the tiny-batch test proves that this model, objective and update path can fit these particular examples. It does not prove that the examples encode the task you intended.
Which is exactly why it sits at the end of the chain rather than the beginning. The task contract is what rules that out, and nothing downstream of it can.
Failing it does not identify a cause either:
it does not tell you whether the problem is capacity, the objective,
the graph, ownership, the update scale or the regularization
It tells you that something in the local system is still wrong and that generalization is not yet the question. The chain above it is what localizes it.
Regularization is a confounder during forensics
If the tiny-batch test fails, the first move is not to change the model. It is to remove the causal paths that were added deliberately. Here is the full recipe on 32 examples the model can memorize perfectly: dropout 0.5, weight decay 0.1, input noise 0.5, label smoothing 0.1.
step 0 train loss 0.74043 eval loss 0.69209 eval acc 0.5000
step 100 train loss 0.41364 eval loss 0.26129 eval acc 0.9375
step 500 train loss 0.37314 eval loss 0.25478 eval acc 0.9688
step 1000 train loss 0.45278 eval loss 0.26875 eval acc 0.9688
step 2000 train loss 0.40803 eval loss 0.24468 eval acc 0.9688
Two thousand steps on 32 examples and it will not go below 0.24. Nothing is broken. Remove the confounders one at a time:
no input noise
step 2000 train loss 0.25678 eval loss 0.07379 eval acc 1.0000
no noise, no smoothing
step 2000 train loss 0.08353 eval loss 0.01601 eval acc 1.0000
+ weight_decay 0
step 2000 train loss 0.02566 eval loss 0.00067 eval acc 1.0000
+ dropout 0
step 2000 train loss 0.00000 eval loss 0.00000 eval acc 1.0000
Each of those is an intervention that removes exactly one source of variation, and every one of them moves the number. None of these mechanisms is bad. Each adds a causal path between the parameters and the loss, and during forensics every extra path is a confounder.
Build the smallest training system that should obviously learn. Then add back one mechanism at a time.
While simplifying, get the mode right, and be precise about what it does. model.train() and model.eval() set the module hierarchy’s training-mode flag. Modules such as dropout and batch normalization — and any custom module that consults self.training — may change behavior in response. These calls do not enable or disable autograd:
train mode, two identical calls agree: False
eval mode, two identical calls agree: True
in eval mode:
loss.requires_grad True
loss.grad_fn NllLossBackward0
gradients present 4 / 4
in train mode inside torch.no_grad():
loss.requires_grad False
loss.grad_fn None
eval() does not disable autograd; a backward pass in eval mode populates every gradient. torch.no_grad() does not disable dropout. They solve different problems and neither substitutes for the other.
Escalation: what to inspect when every boundary passed
Everything so far is cheap — one batch, one step. If all of it is clean and the tiny-batch test still fails, the next instruments cost more and are worth reaching for in this order.
Activation statistics
Forward hooks observe module boundaries. That is a real limitation worth stating: a hook on nn.Linear sees that module’s output, not arbitrary tensor operations written inline in a forward method, so a functional-style model is largely invisible to them.
def activation_report(model, x, types=(nn.Linear, nn.ReLU)):
rows, handles = [], []
def make(name):
def hook(module, inputs, output):
if not torch.is_tensor(output):
return
o = output.detach()
rows.append((name, type(module).__name__, tuple(o.shape),
bool(torch.isfinite(o).all()),
o.mean().item(), o.std().item(),
o.min().item(), o.max().item(),
(o == 0).float().mean().item(),
o.std(dim=0).mean().item()))
return hook
for name, module in model.named_modules():
if isinstance(module, types):
handles.append(module.register_forward_hook(make(name)))
try:
with torch.no_grad():
model(x)
finally:
for h in handles:
h.remove()
return rows
It returns the rows rather than printing them; the formatting below is a separate concern. The try/finally is not decoration. A hook left attached keeps running on every subsequent forward pass, including the ones in your real training loop.
HEALTHY (untrained)
module type shape fin mean std min max zero% batch std
features.0 Linear (256, 32) True -0.019 0.725 -2.675 3.478 0.0% 0.5386
features.1 ReLU (256, 32) True 0.280 0.430 0.000 3.478 51.8% 0.3051
features.2 Linear (256, 32) True -0.059 0.313 -2.263 1.246 0.0% 0.2134
features.3 ReLU (256, 32) True 0.090 0.137 0.000 1.246 52.9% 0.0930
head Linear (256, 2) True 0.104 0.046 0.022 0.353 0.0% 0.0449
HEALTHY (after 300 steps of the reference run)
module type shape fin mean std min max zero% batch std
features.0 Linear (256, 32) True 0.141 1.097 -6.173 5.953 0.0% 0.9362
features.1 ReLU (256, 32) True 0.498 0.700 0.000 5.953 44.5% 0.5866
features.2 Linear (256, 32) True 0.395 5.769 -56.815 46.287 0.0% 4.8293
features.3 ReLU (256, 32) True 2.152 3.823 0.000 46.287 53.8% 3.0047
head Linear (256, 2) True 0.098 54.267 -211.192 208.711 0.0% 53.3090
DEAD FIRST LAYER (features.0.bias = -100)
module type shape fin mean std min max zero% batch std
features.0 Linear (256, 32) True -99.986 0.571 -102.768 -96.809 0.0% 0.5386
features.1 ReLU (256, 32) True 0.000 0.000 0.000 0.000 100.0% 0.0000
features.2 Linear (256, 32) True -0.014 0.116 -0.170 0.176 0.0% 0.0000
features.3 ReLU (256, 32) True 0.044 0.063 0.000 0.176 53.1% 0.0000
head Linear (256, 2) True 0.102 0.067 0.035 0.168 0.0% 0.0000
The zero% column is a trap and this table is the proof. A ReLU is designed to produce zeros: the healthy model reads 51.8% and 52.9%, and after 300 steps of successful training it reads 44.5% and 53.8%. In the dead model, features.3 reads 53.1% — completely normal — while the network has already collapsed two layers earlier. There is no useful “dead ReLU percentage” threshold.
The column that identifies the failure is the last one: standard deviation across the batch, which asks whether the layer’s output depends on which example went in. It reads 0.0000 from features.1 onward. The logits are literally identical for all 256 examples, which is a model that cannot discriminate by construction, and it is a much stronger statement than any zero fraction.
Update scale, which is not gradient scale
An absurd learning rate is worth executing because the trace contradicts the story people tell about it:
step loss grad norm param norm delta norm logits finite
0 6.9416e-01 3.4720e-01 5.3009e+00 3.4720e+03 True
1 9.7960e+07 5.9419e+05 3.4720e+03 5.9419e+09 True
2 1.2313e+26 5.3972e+17 5.9419e+09 inf True
3 nan nan inf nan False
At step 0 the gradient norm is 3.47e-01 — the same value the healthy reference reports, because it is the same model on the same batch. Nothing about the gradient is large. What is large is lr * grad: a parameter tensor with norm 5.3 receives an update with norm 3472. The explosion at steps 1 and 2 is a consequence of the first update, not its cause.
“The gradients exploded” and “the update destroyed the parameters” are different diagnoses. Report gradient norm, parameter norm and delta norm as separate columns, because the first step distinguishes them and every later step does not.
The first non-finite boundary
Do not debug the final NaN. Find the first tensor that became non-finite. Continue the run above and look at the step where the loss first reports nan:
parameters entering step 3:
features.0.weight finite=True max|.|=7.6892e+20
features.0.bias finite=True max|.|=1.2599e+21
features.2.weight finite=True max|.|=2.0724e+21
features.2.bias finite=True max|.|=7.5346e+12
head.weight finite=True max|.|=8.0718e+20
head.bias finite=True max|.|=1.5646e+03
input finite=True inf= 0 max|.|=4.1015e+00
features.0 finite=True inf= 0 max|.|=4.3012e+21
features.1 finite=True inf= 0 max|.|=2.2211e+21
features.2 finite=False inf= 783 max|.|=inf
features.3 finite=True inf= 0 max|.|=4.8522e+34
head finite=False inf= 208 max|.|=inf
first non-finite stage: features.2
Every parameter is finite. The first non-finite value appears inside the forward pass, at features.2, where a matrix of magnitude 2e21 meets activations of magnitude 2e21 and overflows float32.
Then look at features.3. It is finite again. All 783 infinities were negative, and ReLU(-inf) is 0 — the ReLU swallowed the evidence, and it reappears two stages later in the logits. A check placed only on the loss would have reported a NaN three stages downstream of the first overflow, which is itself three steps downstream of the update that caused it. Six boundaries between the symptom and the cause, and every one of them is cheap to look at once you decide to look at the first one instead of the last.
def assert_finite(name, t):
if not torch.isfinite(t).all():
raise RuntimeError(f"{name} first became non-finite")
Explicit staged checks like that are usually enough for forward-pass problems. When the non-finite value first appears during backward() rather than forward(), Chapter 3 covers torch.autograd.detect_anomaly(), which associates the failing backward computation with the forward operation that created it. It is a debugging tool with real overhead; do not leave it enabled.
The learning rate, and what the optimizer actually holds
Now — after task, objective, dependency, gradient, ownership and movement all have evidence — a learning-rate comparison means something. The rule is that every candidate starts from the same model state, the same batch and the same number of steps:
base = build(42)
state = copy.deepcopy(base.state_dict())
for lr in [1e-5, 1e-4, 1e-3, 1e-2, 1e-1, 1.0, 10.0]:
m = TinyClassifier()
m.load_state_dict(state)
...
lr loss @0 loss @100 acc @100 total delta
1e-05 0.6942 0.6940 0.4883 3.4794e-04
1e-04 0.6942 0.6930 0.5156 3.4668e-03
1e-03 0.6942 0.6825 0.7812 3.4065e-02
1e-02 0.6942 0.5925 0.7227 3.1063e-01
1e-01 0.6942 0.0949 1.0000 2.1698e+00
1e+00 0.6942 0.0162 1.0000 4.3333e+00
1e+01 0.6942 1.4909 0.5039 2.2257e+02
The identical loss @0 in every row is the evidence that this is a controlled comparison. Comparing different random initializations and calling it a learning-rate experiment produces a table that looks exactly like this one and means nothing. Chapter 14 owns the general question of when two runs are comparable at all; here the requirement is narrow — same state, same data, same steps.
The total delta column is the interpretable one. At lr=1e-5 the entire model moved by 3.5e-04 in a hundred steps. That is a training run that is not stuck; it is a training run that is too slow to see, and the two require different responses.
Finally, do not read the learning rate off the configuration file, and do not assume param_groups[0] speaks for the optimizer:
group 0: lr=0.0001 weight_decay=0 params=4
group 1: lr=0.1 weight_decay=0.01 params=2
after step 0: lrs = [1e-05, 0.010000000000000002]
after step 1: lrs = [1.0000000000000002e-06, 0.0010000000000000002]
after step 2: lrs = [1.0000000000000002e-07, 0.00010000000000000003]
Two groups with a thousand-fold difference in learning rate and different weight decay, and a StepLR stepped once per batch instead of once per epoch has taken the feature learning rate from 1e-4 to 1e-7 in three steps. After the first scheduler step, param_groups[0]["lr"] correctly reports 1e-05 for group 0. The mistake is treating that one value as the optimizer’s learning rate: it says nothing about the head group, which is simultaneously at 0.01.
Inspect the optimizer’s current state, group by group. Not the learning rate you remember configuring.
Loss, accuracy and what counts as learning
One more measurement trap, because it sends people down the wrong path regularly. Watch the reference model under plain SGD for forty steps:
step loss accuracy mean margin
0 0.69416 0.4766 -0.00160
5 0.68821 0.7266 0.01023
10 0.68253 0.7773 0.02178
20 0.67173 0.7109 0.04443
30 0.66136 0.6836 0.06704
40 0.65142 0.6680 0.08948
Loss falls monotonically. The mean margin between the correct and incorrect logit rises monotonically. Accuracy goes up, then down, from 0.7773 to 0.6680. Nothing is wrong. Accuracy is a step function of the argmax, so it changes only when an example crosses the decision boundary, and early in training a model can improve its average confidence while a handful of examples move the wrong way.
Loss and accuracy answer different questions. A flat accuracy is not evidence of a flat model, and a falling accuracy over a few steps is not evidence of a broken one. Read the loss and a task-appropriate metric together.
The same caution applies to individual steps. Do not assert that every optimizer.step() must reduce the current batch’s loss — a stochastic or adaptive optimizer can take steps that do not, for reasons that have nothing to do with a bug. The one-step invariant established in this chapter is parameter movement under a controlled update; useful loss reduction is a multi-step capability claim. Keep them separate.
Where this chapter stops
Two boundaries are deliberately left closed.
If a model trains correctly in full precision and fails only under mixed precision, this chapter’s job is already done: establish that the same controlled problem learns in float32, then hand the mixed-precision execution path to Chapter 12, which owns autocast, gradient scaling and their interaction with performance. This environment has no CUDA device, so no claim is made about them here.
If the model learns but a new run is worse than an old one, that is a different question with a different method. Chapter 11 uses fixed seeds, a fixed batch and a frozen initial state because a failure that cannot be reproduced cannot be investigated. It does not follow that two full training runs are comparable, and Chapter 14 owns that: run-to-run variation, what changed in the configuration or environment, and whether a validation difference is a regression or ordinary run-to-run variation.
When a failure is intermittent, preserve the evidence rather than trying to reproduce it from memory:
if not torch.isfinite(loss):
torch.save({"x": x_batch.detach().cpu(), "y": y_batch.detach().cpu(),
"model": model.state_dict(),
"optimizer": optimizer.state_dict(), "step": step},
"failing_batch.pt")
raise RuntimeError("non-finite loss")
A saved batch converts a rare failure into a deterministic test case, which is the state in which every technique in this chapter applies.
Using AI on a model that will not learn
An assistant asked “why is my model not learning” will produce a plausible list, because a plausible list is what the question invites. Every item on it will be a real cause of a real failure somewhere. None of them will be evidence about your system.
The prompt below works because it refuses the repair and asks for the localization. It is long deliberately: the structure is the value.
My PyTorch model runs but does not learn.
Do not rewrite the model and do not suggest hyperparameters yet.
I will provide evidence from one fixed batch.
Build a learning-chain table with these boundaries:
1. TASK
- what does one input represent?
- what does its target mean?
- what evidence says they are still aligned?
2. OBJECTIVE
- logits shape and range
- exact loss function and target representation
- manual one-example verification
3. DEPENDENCY
- loss.requires_grad and grad_fn
- the full list of parameters I expect to be trainable
4. GRADIENT
- for each expected parameter:
grad None / exactly zero / finite / non-finite, and grad norm
5. OPTIMIZER OWNERSHIP
- trainable model parameter ids
- optimizer parameter ids
- model parameters missing from the optimizer
- optimizer parameters no longer in the model
6. UPDATE
- parameter norm before
- parameter delta after one step
- the current learning rate for EVERY parameter group
7. CAPABILITY
- tiny fixed-batch initial loss
- loss after controlled repeated steps with the REAL optimizer
- final tiny-batch accuracy
Identify the FIRST boundary whose evidence contradicts the intended
training system.
For that boundary:
- give me the two or three most plausible mechanisms;
- propose the smallest experiment that distinguishes them;
- say which measurement should change if the diagnosis is correct;
- say which measurement must stay unchanged.
Do not propose architecture changes, learning-rate tuning, gradient
clipping, mixed precision, more data or another optimizer until every
earlier boundary has passed.
Three clauses earn their place. Asking for the first boundary rather than a list prevents a plausible downstream fix from being applied to an upstream cause — the UPDATE row failed in two of this chapter’s four ledger runs, with nearly identical text and two completely different causes. Asking what must stay unchanged is what makes the proposed experiment falsifiable rather than merely confirmatory. And asking for the tiny-batch test with the real optimizer closes the loophole that a fresh optimizer silently repairs the exact bug being investigated.
The last paragraph is the one to keep. An assistant given a symptom will reach for the interventions at the end of the chain, because those are the ones most discussed in its training data. Learning rate, optimizer choice, clipping and architecture are all downstream of six boundaries that are cheaper to check and more likely to be broken.
Ask AI to locate the first missing consequence before asking it to repair training.
The learning-chain debugging sequence
1. Freeze one known batch and a known initial state. Make the failure
repeatable before trying to explain it.
2. Prove the TASK. Does each target still describe its input? Use a probe
you built on purpose; no tensor statistic can see this.
3. Prove the OBJECTIVE. Are the model's outputs in the form this loss
expects? Reproduce the loss by hand for one example.
4. Prove the DEPENDENCY. Which of the parameters you intend to train does
this loss actually reach? loss.grad_fn is not that evidence.
5. Prove the GRADIENT. Are the expected gradients present and finite?
Distinguish None from exactly zero from non-finite, and note where in
the model each occurs.
6. Prove OWNERSHIP. Does the optimizer hold those exact parameter objects?
Check both directions: missing from the optimizer, and stale in it.
7. Prove the UPDATE. Snapshot, step, measure deltas. Read the current
learning rate and weight decay for every parameter group.
8. If the update semantics are still unclear, reproduce one step on a copy
with fresh SGD (momentum 0, weight decay 0) and verify d(theta) = -lr * grad.
9. Prove CAPABILITY. Strip augmentation, dropout, weight decay, smoothing
and schedulers, then overfit a tiny fixed batch — with the REAL
optimizer, not a fresh one.
10. If that still fails, escalate: per-layer activation statistics including
across-batch variation, gradient versus update scale, and the first
non-finite stage.
11. Only once the controlled system learns should you reintroduce
regularization, augmentation, schedulers, mixed precision and the full
dataset — one at a time.
12. Change one thing. Rerun from the first boundary that change could affect.
Steps 2 through 4 catch what no gradient inspection can. Steps 5 through 7 catch the opening failure. Step 9 catches what nothing before it can, and proves less than people think it does.
What you should now be able to answer
Here is a ledger-style summary of the kind you will be handed. It is a constructed scenario, not an experiment from this chapter. Work through it before reading on.
loss 0.6931 at step 0, 0.6929 at step 500
accuracy 0.507 throughout
all 14 parameters: requires_grad True, grad present, grad finite
total grad norm 4.1e-06
optimizer: AdamW, 1 group, lr 3e-4, 14 params, 0 missing, 0 stale
all 14 parameters: nonzero delta after one step
tiny batch (32 examples, real optimizer, 2000 steps): 0.6931 -> 0.6902
Which boundary is the first one without support? None of A through F. Task, objective, dependency, gradient, ownership and update all report evidence consistent with a healthy system. The first unsupported link is capability: the model cannot fit 32 examples. That is where the investigation continues, and everything above it has been ruled out cheaply.
Is 4.1e-06 a small gradient norm? Unanswerable as stated, and this is the trap. There is no architecture-independent scale. It becomes evidence when compared: against the same measurement on a known-good run, against the norms at other layers, or against its own trajectory over steps.
The loss sits at 0.6931 and accuracy at 0.507. What does that tell you? That the model is producing near-uniform predictions on a balanced two-class problem, since ln(2) = 0.6931 and the majority baseline is 0.507. It is a precise description of the symptom and says nothing about the cause. A stalled optimizer, a broken graph, a misaligned target relation and a learning rate three orders of magnitude too small all pass through this exact state.
Every parameter has a nonzero delta. Does that prove the optimizer is working correctly? No. It proves the objects moved. Under AdamW the delta ranged from 18 to 212 times lr * ||grad|| in this chapter’s measurements, and decoupled weight decay contributes movement that does not come from the gradient. To check that a specific gradient produces a specific update, use a copy and a fresh plain SGD.
The tiny-batch test passes when you rebuild the optimizer. Is the system fixed? No — you have changed the system under test. That is precisely how the opening failure hides: 0.68758 unchanged with the real optimizer, 0.148 with a fresh one. The gap between those two numbers is the diagnosis.
Training loss falls and validation loss does not. Is that this chapter’s problem? No. Every boundary here is about whether the mechanism works, and it does. A train/validation gap is a question about generalization and about whether the two measurements are comparable, which belongs with Chapters 7 and 14.
The loss plateaus near 0.3133 and accuracy is 100%. What happened? In this chapter’s two-class experiment, pre-softmaxing before cross entropy produces exactly that limiting floor, so it is a strong clue to inspect the objective boundary. It is not a unique fingerprint in an unfamiliar system: verify what tensor is actually being passed to the loss before naming the cause.
A NaN appears in the loss at step 40. Where do you look? Not at step 40’s loss. Trace the forward pass stage by stage for the first tensor that is non-finite, then check whether the parameters entering that step were already enormous, then check the update scale on the step before. This chapter’s example had three stages and three steps between the cause and the NaN — and a ReLU that zeroed the infinities in between, so the non-finite value briefly disappeared and then came back.
What single piece of evidence separates “not learning” from “learning too slowly”? There is no universal single measurement. Parameter movement separates a stationary model from one that is changing, and a controlled loss trajectory tells you whether that movement is useful. In this chapter’s lr=1e-5 run, the model moved by only 3.5e-04 over 100 steps while every mechanical boundary passed and the loss changed only slightly; together those observations support “learning too slowly.” A zero delta localizes a different kind of failure, but nonzero movement alone does not prove useful learning.
Exercises
Gradient but no update. Reproduce the opening: freeze the feature extractor, construct the optimizer, then replace the head. Prove that the new head appears in
named_parameters(), requires gradients, receives a finite gradient, is absent from the optimizer, and has a parameter delta of exactly zero. Then measure how its.gradbehaves over five iterations and explain whyoptimizer.zero_grad()did not clear it. Repair only the optimizer construction.Chapter 1 at model scale. On one fixed batch, using a copy of the model and a fresh
SGD(lr, momentum=0, weight_decay=0), verify that every parameter satisfiesΔθ = -lr · gradto floating-point tolerance. Then repeat withmomentum=0.9, and withAdamW, and record the ratio of||Δθ||tolr·||grad||for each parameter. Explain why the equality is not expected to survive either change.Detach one edge. Insert
detach()between two layers. Identify the first parameter whose gradient becomesNone, and confirm thatloss.requires_gradandloss.grad_fnare both healthy throughout. Then measure the final loss against the reference run and argue why this failure is harder to notice than a complete stop.Zero is not
None. Build a controlled dead-ReLU case. For each parameter, record whether.gradisNone, an exactly-zero tensor, or nonzero, and explain each result from the local derivative. Then explain why one bias in the dead region still receives a gradient.Destroy the alignment. Permute the inputs without permuting the labels. Show that shape, dtype, range, mean, standard deviation and class counts are all unchanged, that the known-rule probe is the only check that fails, and that the full training system remains mechanically healthy. Report held-out accuracy on the real task and compare it with the trivial baseline.
Wrong objective, right mechanism. Compare cross entropy on logits with cross entropy on already-softmaxed probabilities from an identical initialization. Report the loss trajectory, the step-0 gradient norm and the final training accuracy for both. Derive the floor of the broken objective analytically for
Cclasses, and verify it numerically forC = 2andC = 10.Two tiny-batch tests. Run the capability test on a broken system twice: once with the real optimizer and once with a fresh one. Explain why the two answers differ and which question each one answers. Then run it on labels drawn from
randintand report how many steps memorization takes.Simplify one mechanism at a time. Take a model with dropout, weight decay, label smoothing and input noise that will not overfit 32 examples. Remove one mechanism per run, always from the same initial state, and record the final loss. Identify which single mechanism accounts for most of the gap, and state what that does and does not tell you about the full training recipe.
One-batch forensics on an unfamiliar model. Run the learning ledger against a small model you did not write. Identify the first unsupported link before changing anything. Then deliberately introduce a second, later failure, rerun, and confirm the ledger still reports the earlier one.
Next: it learns, but is it using the machine?
Learning is no longer a thing that either happens or does not. It is a chain of observable consequences, and every link leaves evidence on a single batch:
TASK the target still describes the input known-rule probe
OBJECTIVE the scalar is the loss we intended one-example derivation
DEPENDENCY the loss reaches the intended parameters grad present, not None
GRADIENT those parameters received real derivatives None / zero / finite
OWNERSHIP the optimizer holds those exact objects two-way id() audit
UPDATE those exact objects moved snapshot and delta
CAPABILITY repeated controlled steps fit a tiny batch real optimizer, no regularization
Five deliberate failures stressed different parts of that chain, but they were not all discoverable by the same instrument. The stale head satisfied everything up to ownership and never moved. The detached layer reached the gradient boundary, where four expected gradients disappeared even though the fixed feature representation still supported partial learning. The shuffled targets failed at the task boundary and then optimized nonsense competently. The pre-softmaxed objective passed the generic mechanical ledger because deciding whether a differentiable scalar is the intended objective requires an objective-specific check. The absurd learning rate also passed the one-step mechanical ledger; its failure appeared in update scale and only later as non-finite values. That division of labor is the point: the chain tells you which evidence to demand, not that one generic checker can certify every link.
Not one of them raised an exception.
Learning is a chain of observable consequences. A correct target should produce the intended loss; that loss should reach the intended parameters; those parameters should receive finite gradients; the optimizer should own those exact objects; the step should move them; and repeated controlled steps should change the model’s behavior. Find the first consequence that fails, and debug there.
We can now establish that a known sample means what its target says, that the intended objective depends on the intended parameters, that gradients reach them, that the optimizer owns those exact objects, and that those objects move. On a controlled copy with plain SGD we can go further and verify the update equation itself. Finally, repeated controlled steps can establish that the training system is capable of fitting a small batch.
That still says nothing about whether the machine is being used well. A training loop can be entirely correct and waste most of a GPU. It can fit in memory and spend half its time waiting on synchronization. It can be slower after compilation than before it. None of the instruments in this chapter would notice, because all of them measure whether the right thing happened and none of them measures what it cost.
The next chapter starts from a model whose learning we have proved, and asks a completely different question: where is the time going, where is the memory going, and why is the device waiting? Its governing idea is as simple as this chapter’s and just as easy to skip:
Measure first. Optimize second. Measure again.