← PyTorch From First Principles

Assembly: A Language Model You Can Interrogate

Assemble a small GPT-style language model and interrogate it with the book's full diagnostic stack—tokenizer and target contracts, attention interventions, parameter accounting, learning and performance ledgers, compilation, generation, and paired evaluation—so the finished model remains inspectable rather than magical.

Here are two training runs of the same small language model on the same data. The only difference is one line in the batching function.

CORRECT     y = data[i+1 : i+block+1]        the next-token shift
MISALIGNED  y = data[i   : i+block]          no shift

step      correct val    misaligned val
   0          3.6695          2.6469
 250          1.1038          0.0007
 500          0.4730          0.0004
 750          0.3940          0.0003
1000          0.3690          0.0002
1250          0.3581          0.0002
1500          0.3558          0.0001

The misaligned run’s loss falls to 0.0001 — three thousand times lower than the correct run’s 0.36. By every number on the dashboard it is the best training run you have ever seen. Nothing raised. The shapes are identical: x is [B, T], y is [B, T], the loss is a finite scalar with a grad_fn.

Then you generate from each:

correct:    'The planet measures the ancient planet because measures it near the water.
             A sudden market builds a hidden market again. When that planet counts slowly, '
misaligned: 'The                                                                          '

The misaligned model learned the identity task. With y[:, t] == x[:, t], the answer is already present at the same position: the residual stream contains the embedding of token t, and causal self-attention is also allowed to attend to that position itself. The network therefore has an extremely cheap path from the visible token to the target token. The loss rewarded that wrong task the whole way down. What catches it is not a better loss curve — it is one line, checked before training starts:

assert torch.equal(x[:, 1:], y[:, :-1])      # does the target still describe the input?

That is Chapter 11’s first link, and it is the entire chapter in miniature.

The loss can be actively misleading. A metric that improves is not evidence that the model is learning the thing you meant — and a model large enough to hide its mistakes needs every instrument you have.

Where we are

Fourteen chapters, fourteen instruments. Chapter 2 asked which tensor first became wrong rather than first became illegal. Chapter 3 asked where the gradient path stops existing. Chapter 5 asked what PyTorch thinks belongs to your model. Chapter 6 asked where the training loop is actually waiting. Chapter 7 asked what the model actually sees. Chapter 8 derived convolutional geometry into a shape ledger. Chapter 9 made feature space observable. Chapter 10 asked which position is comparing with which. Chapter 11 proved the learning chain one link at a time. Chapter 12 localized performance to a phase and an operator. Chapter 13 found the first compiler assumption that stopped holding. Chapter 14 asked whether a difference between two runs is real.

Each of those was earned from one constructed failure against one small system.

This chapter builds the first system too large to hold in your head — a GPT-style language model, about eight hundred thousand parameters, forty-odd tensors, a forward pass that goes text → tokens → embeddings → attention → residual blocks → logits → loss, and a second execution mode for generation. No single instrument certifies it. But every instrument the book built is now a question you ask of it:

tokenizer + batching   does the target still describe the input?      Chapter 11
embeddings             what does PyTorch think belongs to the model?  Chapter 5
attention              can a future token change an earlier output?   Chapter 10
block assembly         where does the shape ledger first diverge?     Chapter 8
first training steps    which link in the learning chain is broken?    Chapter 11
throughput and memory  which phase, at which sequence length?         Chapter 12
compilation            did it help, and did correctness hold?         Chapter 13
comparison             is this run really different from that one?     Chapter 14

The reader who finishes this chapter should be able to say which instrument answers which question about a transformer, and reach for the right one without being told. That is the book’s real deliverable. The model is the demonstration.

The environment

Every number, trace and generated sample in this chapter came from running the committed code under experiments/ch15/:

Python   3.11.4
PyTorch  2.6.0+cu118
GPU      NVIDIA GeForce RTX 2060, compute capability 7.5

The model is deliberately small: n_layer=4, n_head=4, n_embd=128, block_size=64, character-level vocabulary. It has 804,224 parameters and trains to convergence in under a minute on this GPU; the whole experiment suite, including the seed sweeps, is about twenty minutes of compute. A larger model would not teach anything the instruments do not already show at this size.

The data is a self-contained synthetic corpus — 200,000 characters of English-like sentences generated from a small fixed word list and five templates:

That garden repeats that small garden because repeats it by the door. Near the water,
that hidden shadow carries that shadow. When a river holds again, a golden river holds.

There is no download and no copyrighted text: the corpus is reproducible from a seed, and it has real learnable structure — word boundaries, template grammar, common bigrams, punctuation — so a trained model produces recognisable (if nonsensical) output, which is what makes the generation experiments meaningful. experiments/ch15/corpus.py builds it.

The pipeline

text
  ↓  tokenizer            characters ↔ integer IDs
token IDs  [B, T]
  ↓  batching             x = window;  y = window shifted by one
  ↓  token + position embedding
[B, T, C]
  ↓  transformer blocks   pre-norm attention + MLP, residual, ×n_layer
  ↓  final LayerNorm
  ↓  lm_head              [B, T, vocab] logits
  ↓  cross entropy        vs the shifted targets
loss  → backward → AdamW → repeat
  ↓
generation               feed the model its own output, one token at a time

Two boundaries in that pipeline are genuinely new and get most of this chapter’s attention: text → token IDs, and training → generation. Everything between them is composition of parts the book already built.

Data: the tokenizer

A character tokenizer is a pair of dictionaries. What matters is that it round-trips and that its IDs are in range.

class CharTokenizer:
    def __init__(self, text):
        self.chars = sorted(set(text))
        self.stoi = {c: i for i, c in enumerate(self.chars)}
        self.itos = {i: c for c, i in self.stoi.items()}
        self.vocab_size = len(self.chars)

    def encode(self, s):  return [self.stoi[c] for c in s]
    def decode(self, ids): return "".join(self.itos[int(i)] for i in ids)
  round-trip decode(encode(s)) == s        True   for the full corpus
  full corpus tensor: dtype int64  min 0  max 36  vocab 37

The round-trip check is Chapter 7’s contract idea at the input boundary: prove the transform is lossless before trusting anything downstream. The dtype check matters because token IDs are indices, not continuous values. nn.Embedding expects an integer index tensor; passing floating-point IDs is a boundary-contract error and should be diagnosed there, before the model’s representation is discussed.

The failure the tokenizer causes when it is wrong is an out-of-range ID reaching the embedding table:

emb = nn.Embedding(vocab_size, 8)
emb(torch.tensor([[0, 1, vocab_size]]))       # one past the end
IndexError: index out of range in self

That is the exact error people paste into a search box. The mechanism is a contract mismatch between the IDs produced by the data/tokenizer side and the rows the embedding table actually owns: a tokenizer fit on a different corpus, a stale checkpoint with a smaller vocabulary, or a special-token ID the model was never sized for can all create it. A tokenizer fit on a 2,000-character sample of this corpus, for instance, is missing the character S and raises KeyError: 'S' the moment it meets a capital-S sentence in the full text. Verify the tokenizer round trip, the ID range, and the model’s vocabulary size at the boundary, and this class of failure never reaches training.

Data: batching and the alignment contract

def get_batch(source, batch_size, block_size):
    starts = torch.randint(0, len(source) - block_size - 1, (batch_size,))
    x = torch.stack([source[i : i + block_size]     for i in starts])
    y = torch.stack([source[i + 1 : i + block_size + 1] for i in starts])
    return x, y

Every position of x predicts the next token, so y is x shifted left by one. The contract:

assert torch.equal(x[:, 1:], y[:, :-1])

This is the check from the opening. It is worth restating why it is the single most valuable line in the build: the misalignment failure is invisible in the loss (the loss improves), invisible in the shapes (both [B, T]), invisible in the gradients (they are finite and flow everywhere), and it produces a model that scores brilliantly and generates nothing. The only place it is visible is here, in the relationship between the input and the target — Chapter 11’s TASK boundary, the first thing to prove and the first thing people skip.

The model: embeddings and positions

self.token_embedding = nn.Embedding(vocab_size, n_embd)      # [V, C]
self.position_embedding = nn.Embedding(block_size, n_embd)   # [block, C]
...
x = self.drop(self.token_embedding(idx) + self.position_embedding(pos))

Content-based self-attention without a positional signal or order-dependent mask is permutation-equivariant. This model’s causal mask already introduces an ordering constraint — earlier positions cannot see later ones — but it does not supply the explicit learned position identity that the model uses here. That comes from the position embedding. idx is [B, T]; token_embedding(idx) is [B, T, C]; position_embedding(arange(T)) is [T, C] and broadcasts across the batch. This is Chapter 2’s broadcasting, used deliberately.

Both embedding tables are nn.Module attributes, so they register — they appear in the model’s registered parameter/state structure and, once the optimizer is built over model.parameters(), its trainable parameters can be owned by the optimizer. Chapter 5’s question — does PyTorch think this belongs to the model? — is answered here by construction. Weight tying later makes object identity important again: sharing two names is not the same thing as owning two independent parameters, and tying after optimizer construction would create exactly the kind of stale-parameter mismatch Chapter 11 taught us to audit.

The model: attention

Chapter 10 built attention carefully: the head split as a partition of the feature axis, the [B, H, T, T] score matrix, the softmax over keys, the causal mask, and — crucially — the four levels at which “correct” can mean four different things (shape, axis semantics, numerical invariants, behavior). This chapter does not re-derive any of that. It packages it:

def split_heads(x, n_head):
    B, T, C = x.shape
    return x.view(B, T, n_head, C // n_head).transpose(1, 2)   # [B, H, T, D]

def merge_heads(x):
    B, H, T, D = x.shape
    return x.transpose(1, 2).contiguous().view(B, T, H * D)     # [B, T, C]

class CausalSelfAttention(nn.Module):
    def forward(self, x):
        q, k, v = self.qkv(x).chunk(3, dim=-1)
        q, k, v = (split_heads(t, self.n_head) for t in (q, k, v))
        y = F.scaled_dot_product_attention(q, k, v, is_causal=True,
                                           dropout_p=self.dropout if self.training else 0.0)
        return self.resid_dropout(self.proj(merge_heads(y)))

F.scaled_dot_product_attention with is_causal=True is the whole mechanism — Chapter 10 verified it against a hand-written implementation. The two things that are still worth checking at model scale are Level 2 (does the head split preserve the data?) and Level 4 (is it actually causal?), and neither is visible in a shape.

Level 2 — the head split round trip. split_heads followed by merge_heads must return the original tensor exactly. The wrong version — q.reshape(B, H, T, D) instead of q.view(B, T, H, D).transpose(1, 2) — produces a tensor of identical shape that has cut the sequence into head-sized blocks instead of partitioning the feature axis. The shape ledger below catches it; a shape assertion does not.

Level 4 — causality by intervention. A triangular mask is a claim. The evidence is behavioral: put the model in evaluation mode, change only the tokens after position i, rerun, and require the outputs at positions 0..i to agree within a stated numerical tolerance while some suffix output actually changes.

On the correct model, after training:

  max |Δ logits| at positions 0..i  :  0.000e+00      the past did not move
  max |Δ logits| at positions i+1.. :  varies         the future did

This run produced exact zero on the prefix; the stronger portable contract is that the prefix delta stays within the tolerance appropriate to the backend and dtype. The second number matters too: if the suffix delta were also zero, the test could be passing because the intervention had no observable effect.

Now set is_causal=False and train the same model:

                        causal      non-causal
  val loss (1000 steps)  0.369        0.023          16x lower -- looks extraordinary
  past Δ after training   0.000e+00    1.03e+01       the past moved by 10 logits

  causal sample:      'The planet measures the ancient planet because measures it by the door.'
  non-causal sample:  'The       h    ,      hh  h   ,,  ,,, , ,    h h h h h h h h h h h h h h'

The non-causal model can attend from position t to position t+1, whose input token is exactly the next-token target for position t (except at the boundary). The loss rewards that leakage all the way down. Inspecting the mask can suggest the bug; the future-token intervention is the behavioral evidence that proves the model actually uses information it should not have — Chapter 10’s Level 4, at full scale.

The model: MLP, block, assembly

The rest is composition. The per-token MLP (Linear → GELU → Linear → Dropout, hidden 4C) operates independently at every position and preserves [B, T, C]. The pre-norm block wraps attention and MLP in residual connections:

def forward(self, x):
    x = x + self.attn(self.ln1(x))
    x = x + self.mlp(self.ln2(x))
    return x

Every block preserves [B, T, C], which is the invariant that makes stacking them trivial. The full model adds the embeddings, the block stack, a final LayerNorm, and the lm_head projection to vocabulary logits.

To verify the assembly, run Chapter 8’s shape ledger through the whole forward pass — a row per stage, each carrying the shape and the invariant that stage owes:

  SHAPE LEDGER   B=2 T=16 C=128 H=4 V=37
    token_embedding    (2, 16, 128)   (B,T,C)=(2,16,128)                  ok
    position_embedding (16, 128)      (T,C)=(16,128)                      ok
    tok + pos          (2, 16, 128)   (B,T,C) by broadcast                ok
    head split         (2, 4, 16, 32) merge(split(q)) == q  (round trip)  ok
    block[0..3]        (2, 16, 128)   (B,T,C) preserved                   ok
    lm_head            (2, 16, 37)    (B,T,V)=(2,16,37)                    ok
    flatten for CE     (32, 37)       (B*T,V)=(32,37)                     ok
  first divergence: none

Now build the model with the wrong head reshape and run the same ledger:

    head split         (2, 4, 16, 32) merge(split(q)) == q  (round trip)  X  <-- FIRST DIVERGENCE
    block[0..3]        (2, 16, 128)   (B,T,C) preserved                   ok
    lm_head            (2, 16, 37)    (B,T,V)=(2,16,37)                    ok
  first divergence: head split

The broken split’s output shape is (2, 4, 16, 32) — identical to the correct one. Every downstream row still reports ok. A forward pass would not raise, and the model would train to a mediocre loss and never say why. The round-trip row is the one that sees it, because it checks the data, not the dimensions. This is Chapter 8’s first-divergence rule and Chapter 10’s Level 2, composed.

Parameter accounting

def param_count(model):
    seen = set()
    total = 0
    for p in model.parameters():
        if id(p) in seen:
            continue
        seen.add(id(p))
        total += p.numel()
    return total
  THIS MODEL   n_layer=4 n_head=4 n_embd=128 vocab=37       total 808,960
    mlp                             526,848   65.1%
    attention                       262,144   32.4%
    embeddings (tok + pos)           12,928    1.6%
    lm_head (output projection)       4,736    0.6%
    layernorm                         2,304    0.3%

There is a widely repeated claim that the embeddings dominate a small language model. At subword scale that is true — the same accounting at GPT-2-small scale:

  n_layer=12 n_embd=768 vocab=50257                          total 124,356,864
    mlp                          56,623,104   45.5%
    embeddings (tok + pos)       39,383,808   31.7%     <- the token table alone is 31%
    attention                    28,311,552   22.8%

— the token embedding alone is about 38.6M parameters, roughly 31% of the canonical 124.4M tied GPT-2-small count. But at character scale the vocabulary is 37, the token table is only 37 × 128 = 4,736 values, and the MLPs dominate. The important variable is the vocabulary-to-model-width/body ratio, not a universal rule that “embeddings dominate.” Knowing which regime you are in tells you how significant the next section can be.

Weight tying

The input embedding maps a token ID to a vector; the output lm_head maps a vector back to a distribution over tokens. Both are [vocab, n_embd] matrices about the same vocabulary, and a standard design shares them:

self.lm_head.weight = self.token_embedding.weight

Verify the tie took — Chapter 5’s discipline, because an assignment that runs is not proof of anything:

  tied:   lm_head.weight is token_embedding.weight  ->  True
  untied: same object                               ->  False
  config   params   parameter tensors   shares storage
  tied     804,224           44          True
  untied   808,960           45          False

At the GPT-2-small configuration above, an untied output projection would add another 50,257 × 768 = 38,597,376 parameters. Tying avoids that matrix: about 31% of the canonical tied model’s parameter count, or about 24% of the hypothetical untied total. Here it removes only 4,736, about 0.6% of the untied model. So the question for this model is not primarily memory; it is whether sharing the input and output representation changes behavior enough to resolve. A paired seed sweep (four seeds, 2,000 steps each, same seed on both sides):

  seed     tied   untied   untied - tied
     0   0.3525   0.3520      -0.0005
     1   0.3517   0.3519      +0.0002
     2   0.3554   0.3527      -0.0027
     3   0.3518   0.3515      -0.0003
  mean untied - tied: -0.0008   stdev 0.0011   (1/4 favour tied)

The mean per-seed difference is -0.0008 (untied - tied) with a standard deviation of 0.0011, and the four paired deltas do not establish a stable advantage in either direction. The six-seed baseline variation below is also much larger than the observed mean shift, but that range is context, not a statistical equivalence threshold. The honest result is not resolved with four paired seeds. The parameter saving is exact; the representational effect is an empirical question, and this experiment is not strong enough to call it better, worse, or equivalent.

Training: prove it is happening

Chapter 11 turned “is the model learning?” into a chain of checkable consequences on one fixed batch:

TASK → OBJECTIVE → DEPENDENCY → GRADIENT → OWNERSHIP → UPDATE → CAPABILITY

Run its learning ledger on the transformer, with two capstone-scale additions. OBJECTIVE gains an initialization sanity check: a uniform predictor has cross-entropy ln(vocab_size), so for this deliberately small-logit initialization an initial loss on the same scale is expected. That is a reference, not a universal invariant for every language-model initialization. And CAPABILITY is overfit one tiny batch, which for a model this size is the strongest single end-to-end proof that the training machinery can fit those examples.

  LEARNING LEDGER   TinyGPT   batch (8, 32)
    TASK      alignment       x[:,1:] == y[:,:-1]                       ok
              id range        [0, 36] of 37                             ok
    OBJECTIVE initial loss    3.668   ln(vocab) = 3.611   ratio 1.02    ok
              one-example CE  manual 3.52942 vs 3.52942                 ok
    GRADIENT  present         44/44                                     ok
              finite          44/44                                     ok
    OWNERSHIP owned           44/44                                     ok
    UPDATE    moved           44/44                                     ok
  first divergence: none

  CAPABILITY -- overfit one tiny batch (8 x 32):
    step   0  loss 3.6557
    step 100  loss 0.0230
    step 200  loss 0.0083
    step 300  loss 0.0045
    step 399  loss 0.0029

Every link holds; the model drives one batch to near-zero loss in 400 steps. Now the fault the initial-loss check exists to catch. Tie the weights but skip the small-scale initialization — leave the embedding table at PyTorch’s default N(0, 1), which the tied lm_head then inherits:

    TASK      alignment       x[:,1:] == y[:,:-1]                       ok
    OBJECTIVE initial loss    119.463   ln(vocab) = 3.611   ratio 33.08  <-- FIRST DIVERGENCE
    GRADIENT  present         44/44                                     ok
    OWNERSHIP owned           44/44                                     ok
    UPDATE    moved           44/44                                     ok
  first divergence: OBJECTIVE/initial loss

The chain passes TASK and first looks wrong at the OBJECTIVE measurement: with the tied output matrix inheriting the embedding table’s N(0, 1) scale, this run produces enormous, spiky logits and a loss thirty-three times the uniform-predictor reference. The downstream mechanical checks still report health — gradients exist, the optimizer owns the parameters, and one step moves them. That proves the system can update from this starting point; it does not prove that training will recover cleanly or quantify how badly it will behave without running that experiment. The ledger has already done its job: it names the earliest surprising consequence before a falling loss can be mistaken for success.

Performance

Chapter 12’s performance ledger: decompose the step, then sweep the axes that matter.

  WORKLOAD  TinyGPT 4L/4H/128d | training step | batch 64 x 64 | CUDA

  PHASE DECOMPOSITION (median ms, synchronized)
    forward       3.428   31.5%
    loss          0.165    1.5%
    backward      6.176   56.8%
    optimizer     1.104   10.2%
    sum          10.872

Backward dominates at 57% — the opposite of Chapter 12’s shallow MLP, where the AdamW step over 30 million parameters was the largest phase. Here the model has only about 800,000 parameters, so the optimizer is comparatively cheap, while backpropagation through the four transformer blocks dominates the measured step. This phase decomposition does not by itself prove that scaled_dot_product_attention is the dominant operator inside backward; that would require the profiler or another operator-level experiment. The Chapter 12 lesson holds: localize before optimizing, and do not localize more finely than the evidence supports.

The batch-size and sequence-length sweeps:

  BATCH SIZE SWEEP (block_size 64)
     batch   ms/step   tokens/s   peak MiB
        16     9.98     102,601      77.1
        32     9.74     210,347     113.6
        64    11.04     370,930     187.2
       128    17.04     480,788     337.1
       256    31.19     525,278     634.6

  SEQUENCE LENGTH SWEEP (batch 32) -- the transformer-specific axis
     block   ms/step   tokens/s   peak MiB
        16    10.09      50,721      56.2
        32    10.14     101,025      77.0
        64     9.72     210,728     113.6
       128    11.47     357,221     187.3
       256    20.65     396,659     337.3

Batch size and sequence length both raise throughput with diminishing returns and both raise peak memory — but not at the same rate. Doubling the batch from 128 to 256 roughly doubles the measured peak; doubling the sequence length from 128 to 256 raises peak memory and step time by about 1.8× in this run. Chapter 12 measured a manual attention implementation that explicitly materialized [B, H, T, T] score/probability tensors and therefore exposed quadratic memory growth directly. The SDPA backend selected in this experiment is memory-efficient and does not materialize those same full tensors, so measured peak allocation grows much more gently over this range. The logical dense attention relation is still pairwise in sequence length; physical allocation depends on the selected kernel.

Compilation

Chapter 13’s question, asked once: did torch.compile help this model, and did correctness hold?

  WORKLOAD  TinyGPT training step | batch 64 x 64 | float32 | CUDA

  config                    first call    steady ms   vs eager
  eager                             -         9.37      1.00x
  compiled                    ~60 s cold      8.12      1.15x
                              ~2 s warm

  correctness: max |Δ logits| eager vs compiled, same weights = 1.9e-06

A modest steady-state speedup — about 1.15× in this measurement, with other runs reaching roughly 1.3× — comes with a first-call cost of about a minute on a cold Inductor cache and about two seconds when the cache is warm. The compiled logits differ from eager by at most 1.9e-06 for the tested input and weights, so the correctness check passes at that stated tolerance; the measurement alone does not need to assign the difference to a particular fusion/reordering mechanism.

The amortization is easy to calculate. The measured steady-state saving is 9.37 - 8.12 = 1.25 ms/step. A 60 s cold compile therefore needs roughly 60,000 / 1.25 ≈ 48,000 reused steps to break even. A 2 s warm-cache first call needs roughly 1,600. The canonical 3,000-step run easily amortizes the warm cost but not the cold one. Whether compilation is worth taking depends on cache state and workload lifetime, not merely on the existence of a steady-state speedup.

That is the whole compiler section. Chapter 13 owns graph breaks, guards, recompilation and the cold-versus-warm cache split; if the compiled model were slower or kept recompiling, the investigation would move there. Here it is one row of the ledger.

Generation is a different program

Training and generation run the same weights through different code. Training is one forward pass over a fixed batch with targets. Generation is a loop that feeds the model its own output, and it has failure modes training never exercises.

Sampling strategy, same trained model, same seed:

  greedy (argmax)      'The machine crosses the machine, because the gentle machine crosses at
                        dawn. That machine carries that machine, because that gentle machine...'
  temperature 0.5      'The hidden market crosses the hidden market again. The golden window
                        counts the golden window without warning. That signal builds that...'
  temperature 1.0      'The quiet shadow remembers the quiet shadow by the door. When one
                        garden watches without warning, one restless garden watches...'
  temperature 1.5      'The quiet small window builds Quietly. Quietly, that restless window
                        answers that window. For years, this narrow forest thout warning...'
  temp 1.0 + top_k 10  'The hidden market crosses the hidden market by the door. When one
                        garden watches without warning, one restless garden watches...'

Greedy is deterministic and, in this sample, falls into a strongly repetitive continuation (“the machine crosses the machine…”). Temperature 0.5 sharpens toward high-probability continuations; temperature 1.5 flattens the distribution enough that this sample contains misspellings (thout warning) and odd capitalization; top_k removes all but the highest-ranked candidates before sampling. None of these settings is universally “correct” — they alter the sampling distribution, and the useful tradeoff depends on the generation task.

Dropout left active during generation. model.eval() is not optional here — it is what turns dropout off. Generate with the model in train() mode and every call gives a different answer, because dropout’s stochasticity is layered on top of the sampling’s. This is Chapter 11’s eval()-versus-no_grad() distinction arriving as a generation bug: torch.no_grad() saves memory but does not touch dropout; model.eval() is the one that matters. TinyGPT.generate calls self.eval() and restores the previous mode in a finally — the pattern from Chapter 10’s attention module.

Generating past the context length. The model has a block_size-row position table. Feed it a longer sequence and it raises at the contract:

  ValueError: sequence length 84 exceeds block_size 64

The generate loop truncates with idx[:, -block_size:] every step; delete that line and the context eventually grows past block_size and hits this. The explicit raise in forward is Chapter 8’s principle — fail at the contract with a message that names the problem, not forty lines later inside an embedding lookup.

The cost structure. Naive generation re-encodes the entire visible prefix at every step:

  new tokens   wall s   ms/token
          32    0.102     3.183
          64    0.219     3.418
         128    0.413     3.229
         256    0.863     3.370
         512    1.730     3.379

Milliseconds-per-token is roughly flat here because block_size is 64 and truncation caps the visible context: once the window is full, each call recomputes a transformer forward over at most 64 positions. Below that cap, naive generation reruns the entire visible prefix for every new token. In dense self-attention, that forward contains O(T²) attention interactions at visible length T (plus O(T) per-token projections/MLPs), so the per-generated-token cost grows with context before the window saturates.

A KV cache changes the computation rather than merely making the same loop faster: each layer stores past keys and values, computes the new token’s projections once, and attends the new query over the cached T keys/values, making the attention work for the new token O(T) rather than recomputing a full T × T attention pass. It also introduces mutable generation state, cache memory, eviction/window semantics and new correctness contracts. That is why it is deliberately outside this build: establish the stateless baseline first.

The run record, and this model’s baseline variation

Before comparing any two training runs of this model — tied versus untied, a new learning rate, a refactor — Chapter 14 says to know how much two runs of the identical configuration differ from the seed alone.

  ONE RUN RECORD
  {
    "seed": 0,           "config_hash": "0ead31df3032",   "params": 804224,
    "torch": "2.6.0+cu118",                               "device": "cuda",
    "final_train_loss": 0.3487,   "final_val_loss": 0.3525,
    "peak_mib": 174.9,            "wall_s": 31.54
  }

    BASELINE VARIATION   TinyGPT 4L/4H/128d, 2000 steps, 6 seeds
    seed    train      val
       0   0.3487   0.3525
       1   0.3475   0.3517
       2   0.3505   0.3554
       3   0.3459   0.3518
       4   0.3484   0.3493
       5   0.3476   0.3515
    val loss   median 0.3518   stdev 0.0018   spread 0.0061
    wall time  median 28.8 s    spread 9.6 s   (machine contention, not the model)

Across these six identical-configuration seeds, validation loss spans 0.0061 with a standard deviation of 0.0018. That empirical variation gives scale to the comparisons: the misalignment, causality and initialization failures move the metric enormously relative to it, so their effects are not subtle. Weight tying moves the paired mean by only 0.0008, and four paired seeds do not resolve that effect. Do not promote the six-seed min/max spread into a universal detection threshold — Chapter 14 showed that paired comparisons can sometimes resolve effects smaller than an unpaired range. The honest statement here is narrower: with the paired evidence collected, weight tying remains not resolved.

We have removed most of the magic

Look back at what the model is:

embedding lookup
+ learned positional representation
+ repeated (pre-norm attention, pre-norm MLP) residual blocks
+ final layer normalization
+ linear projection to vocabulary logits
+ cross entropy against the next-token shift
+ gradient descent with AdamW

There is no pipeline() call, no pretrained checkpoint, no hidden service. Every tensor in the forward pass is one you can name, shape, and trace to its source.

That does not make frontier systems simple. Scale changes the engineering entirely: distributed training across thousands of accelerators, data pipelines measured in trillions of tokens, custom attention kernels, learning-rate schedules and optimizer variants, architectural refinements, quantization, and inference systems with their own deep stack. Those layers introduce failure modes that this small model does not contain.

But the core path is no longer opaque. When a larger system does something unexpected, the questions this book taught remain useful — what tensor is this, which contract should hold, where did the first consequence diverge? At scale those questions are joined by new distributed and systems questions, and the experiments become more expensive, but the discipline of demanding evidence survives.

The LLM-era programmer

An assistant can generate every class in this chapter. It can write the tokenizer, the attention module, the training loop, the sampling function, the checkpoint code. It can generate a confident, fluent explanation of why your loss is not decreasing. Typing the code is no longer the valuable part.

Understanding the runtime is. When the generated system misbehaves, the useful questions are:

Does the target still describe the input?            (alignment)
Are the token IDs in range?                          (tokenizer)
Can a future token change an earlier output?          (causality)
Where does the shape ledger first diverge?            (assembly)
Is the initial loss plausible for the declared initialization, with ln(vocab) as the uniform reference? (initialization)
Do gradients exist and are they finite?               (the chain)
Does the optimizer own the parameters?               (ownership)
Did the parameters actually move?                     (update)
Can the model overfit one tiny batch?                 (capability)
Which phase, at which sequence length, is the cost?   (performance)
Did compilation help, and did correctness hold?       (compilation)
Is this run really different from the last one?       (comparison)

Source code can suggest answers to many of those questions, but it cannot establish the runtime facts by inspection alone. The book’s instruments turn the claims into observations: actual shapes, actual gradients, actual parameter identities, actual intervention deltas, actual timings. That is why so much of this book was about debugging: in a workflow where code is generated in seconds, evidence about what that code actually did is the scarce resource.

Using AI on a generated language model

A transformer is exactly the artifact an assistant will most confidently produce and most confidently misdiagnose. The prompt below refuses the explanation and demands the chain.

Here is a generated GPT-style language model that trains but generates poorly
(or trains suspiciously well, or is slow). Do not propose a fix or explain the
loss yet.

Prove the chain, in order, and stop at the first link without evidence:

 1. TOKENIZER   decode(encode(s)) == s on real text; every ID in [0, vocab).
 2. ALIGNMENT   torch.equal(x[:, 1:], y[:, :-1]) on a real batch.
  3. CAUSALITY   put the model in eval mode; change only tokens after position i.
                Report the max delta on the prefix AND suffix. The prefix must
                remain within a stated numerical tolerance and the suffix should
                change, or the intervention is not informative.
 4. SHAPE       a per-stage ledger from [B,T] to logits, with the head-split
                round trip merge(split(q)) == q. Name the first divergence.
 5. INIT        compare initial loss with the initialization contract; use
                ln(vocab_size) as the uniform-predictor reference, not a universal
                required value.
 6. REGISTRATION every expected parameter in named_parameters(); note any tied.
 7. GRADIENTS   present and finite for every parameter after one backward.
 8. OWNERSHIP   the optimizer's parameter ids vs the model's.
 9. UPDATE      snapshot, step, confirm every parameter moved.
10. CAPABILITY  can it overfit one tiny fixed batch to near-zero loss?

Only after all ten hold: is the loss actually good (compare to ln(vocab) and to
a validation set), and where does the step spend its time?

Do not suggest architecture changes, hyperparameter tuning, more data, a better
tokenizer, mixed precision or torch.compile until a specific link has failed.

The clause that earns its place is the third: report the max delta on the prefix and on the suffix. An assistant asked to “check causality” will inspect the mask and pronounce it triangular. The intervention is falsifiable — a non-zero prefix delta is a fact, and a zero suffix delta means the test proved nothing — and it is the check that would have caught this chapter’s most convincing failure.

An assistant can write the whole model and can inspect the batching code, but source inspection is not proof that the runtime targets are aligned. Ask it to prove the chain, not merely to explain the loss.

What you should now be able to answer

“The loss is dropping fast — training is working.” Suspiciously fast? On this corpus the correct model bottoms out near ln(vocab)/10; a loss heading for 0.001 means the model has found something cheaper to predict than the next token. Check torch.equal(x[:, 1:], y[:, :-1]) and check causality by intervention.

“Validation loss is excellent.” Better than the correct model’s? A non-causal model in this chapter reached 0.023 against the causal model’s 0.37. The metric rewards a model that can see the answer. Run the future-token intervention before celebrating.

“The model runs, so the shapes are right.” The wrong head reshape produces the right final shape and every downstream shape. Trace the ledger; the row that catches it is the head-split round trip, not any shape assertion.

“Generation produces garbage, so the model didn’t train.” Or: dropout is still active (model.eval()?), the context is not truncated to block_size, or the sampling distribution has been collapsed or distorted by an extreme temperature / top_k choice. top_k=1 is effectively greedy selection, not evidence of a broken model. Generation is a separate program with separate bugs; check them before re-training.

“It’s slow.” Which phase, at which sequence length, measured how? The performance ledger decomposes the step; the sequence-length sweep is the transformer-specific axis. “Slow” without a phase and a shape is not a diagnosis.

“I’ll add AMP and compile it.” Against which baseline, and did correctness hold? Chapter 12’s contract and Chapter 13’s correctness check apply unchanged. A faster wrong answer is not an optimization.

“The LLM wrote my training loop, and it looks right.” Then prove the chain: alignment, initial loss, gradients, ownership, movement, overfit-one-batch. “Looks right” is the state every failure in this chapter was in.

Exercises

  1. The misalignment failure. Reproduce the opening: train with y = data[i : i+block]. Record both loss curves and a generated sample from each. Then explain, from the causal mask, exactly why predicting token t from token t is nearly free.

  2. Tokenizer forensics. Build a tokenizer on a truncated sample of a corpus, then encode the full corpus with it. Catch the KeyError. Then feed an out-of-range ID to nn.Embedding and catch the IndexError. Write the two-line check that would have prevented both.

  3. Causality by intervention. Take a trained causal model. Change only the tokens after position i and confirm the outputs at 0..i do not move. Then set is_causal=False, retrain, and show the validation loss improving while the intervention fails. Report both deltas.

  4. The shape ledger. Implement the per-stage ledger for the full model. Break the head split with reshape(B, H, T, D), confirm the output shape is unchanged, and show the round-trip row catching it. Then break the merge instead and show the same.

  5. Parameter accounting at two scales. Count this model’s parameters by module. Then compute the same breakdown for a config with vocab_size=32000, n_embd=512. Identify the vocabulary size at which the embedding table overtakes one transformer block, and explain what that means for weight tying.

  6. Weight tying, measured. Run a paired seed sweep of tied versus untied on your workload. Report the per-seed deltas, their mean and spread, and the baseline seed variation for context. State whether the behavioral effect is resolved; do not use the baseline min/max range as an equivalence threshold. Then decide whether you would tie the weights, separating the exact parameter saving from the uncertain behavioral effect.

  7. The learning ledger. Run the chain on a fresh model. Then introduce, one at a time: a detached residual stream, an optimizer built before a submodule is replaced, and a 10x learning rate. For each, name the first link the ledger reports.

  8. Generation cost. Measure milliseconds-per-token as a function of visible context length below block_size, then show what happens after the sliding window saturates. Explain why naive full-prefix dense attention recomputes O(T²) pairwise work per forward before the cap. Then estimate how a KV cache changes the new-token attention work to O(T) while adding cache memory and state.

  9. This model’s baseline variation. Run the identical configuration at eight seeds. Report the validation-loss distribution. Then make a change you believe helps and evaluate it at the same eight seeds. Use the paired deltas to judge the effect, with the baseline distribution as context rather than treating its min/max spread as a formal threshold.

  10. The whole pipeline, on an unfamiliar model. Take a GPT implementation you did not write. Without changing it, run every instrument in this chapter in order and produce a one-line verdict for each: tokenizer, alignment, causality, shape ledger, initialization, learning ledger, performance ledger, generation modes. Identify the first that fails, or certify that none did.

The final checklist

Organised by instrument, and every item is one this chapter demonstrated.

TASK (Chapter 11)
  [ ] decode(encode(text)) == text
  [ ] every token ID in [0, vocab_size)
  [ ] torch.equal(x[:, 1:], y[:, :-1])          the next-token shift
  [ ] train/validation split is a real holdout

MODEL STRUCTURE (Chapters 5, 8, 10)
  [ ] n_embd % n_head == 0
  [ ] split_heads then merge_heads round-trips exactly
  [ ] every block preserves [B, T, C]; logits are [B, T, vocab]
  [ ] every expected parameter is in named_parameters(); tied weights verified
    [ ] future-token intervention leaves the prefix unchanged within tolerance and changes the suffix

OBJECTIVE AND LEARNING (Chapter 11)
    [ ] initial loss is plausible for the declared initialization; ln(vocab_size) is the uniform reference
  [ ] loss is finite; one-example cross entropy reproduces by hand
  [ ] gradients present and finite for every parameter
  [ ] optimizer owns every trainable parameter
  [ ] parameters move after a step
  [ ] the model overfits one tiny batch

PERFORMANCE (Chapters 12, 13)
  [ ] the step is decomposed into phases, synchronized
  [ ] throughput and peak memory measured, not assumed
  [ ] batch-size and sequence-length behaviour measured
  [ ] compilation benchmarked warm, correctness checked

GENERATION (this chapter)
  [ ] greedy works before sampling is added
  [ ] context truncated to block_size every step
  [ ] model.eval() -- dropout is off
  [ ] temperature > 0; top-k bounded by vocab size; probs finite

COMPARABILITY (Chapter 14)
    [ ] this model's seed-only baseline variation is measured
  [ ] comparisons use shared seeds when pairing is valid
  [ ] the run record carries config, tokenizer, seed, environment

What this book built, and what it did not

Fifteen chapters, starting from one scalar parameter and ending with a language model you can take apart. The through-line was never the model — it was the practice of making hidden structure visible and finding the first place it diverges from what you intended.

What the book deliberately left out, and where each now attaches:

  • KV caching and fast inference — attach to this chapter’s generation loop: the naive full-prefix forward recomputes dense attention over the visible context, while a cache changes new-token attention to reuse past keys/values and perform O(T) attention work against them.
  • Distributed training — attaches to Chapter 12’s phase decomposition and Chapter 5’s parameter ownership, now across processes.
  • Better tokenization (BPE, byte-level) — attaches to the text→IDs boundary and Chapter 7’s representation contract; the model does not care how the IDs were made.
  • Quantization and mixed precision beyond AMP — attach to Chapter 12’s dtype and memory accounting.
  • Serving, batching, streaming — attach to Chapter 6’s producer/consumer model and Chapter 12’s latency-versus-throughput distinction.
  • Evaluation at scale — attaches to Chapter 14’s baseline-variation measurement and eval-path fingerprint.

None of those was required to expose the core mechanisms in this first-principles build. At production scale several become foundational engineering concerns in their own right. The advantage is that they now attach to a system you can already inspect, so you can approach them asking what state was added, which contract changed, and what evidence would show whether the intervention helped — rather than as isolated API calls to be pasted in and hoped over.

Final thought

The most useful skill in PyTorch is not remembering function names. It is being able to look at a running model and ask:

What tensor is this?
What shape should it have, and what does each dimension mean?
Where did this value come from?
Does the gradient reach this parameter, and did the optimizer move it?
Is the target still describing the input?
What did this optimization actually improve?
Is this run really better, or does it just look that way?

If you can answer those, the abstractions stop being magic. And when an assistant writes the next thousand lines for you, you still know how to find out whether they work.

That was the whole point. The GPT-style model matters, but it was never the destination. The destination is the moment a model does something you did not expect and you know, without guessing, how to find out why.