Book II · Closing
Foundations, gathered
The whole of Book I is one plate. Text becomes tokens, tokens become vectors, position is added or applied, and then the same block runs some number of times: mix across positions, compute within one, add the result back into the stream. At the end the stream is projected back to the vocabulary and a distribution comes out.
If any stage of that map is one you could not explain to somebody else, the chapter is named beside it.
The propositions were written to be read in order, because they depend on one another in a specific way: a tensor is only interesting once you know a model is a function of parameters; attention only makes sense once you know why a recurrence could not do it; the residual stream only matters once you have seen a long product of Jacobians decay. The dependency graph on each page is the real structure, and it is worth walking backwards from any proposition that does not land.
One thing to carry forward before the next book. Of everything here, the proposition that constrains all the others is the last one. Architecture, loss, schedule and precision are judged by a number, and the number is worth exactly what the split makes it worth.
The equations
A model
ŷ = f(x ; θ)Two arguments. The world supplies the first; training may only change the second.
A linear layer
y = Wx + bAffine, and affine maps compose into a single affine map. Depth without a non-linearity buys nothing.
The gradient step
θ ← θ − η ∇θ LThe gradient gives the direction; η is where the distance has to come from.
Backpropagation
∂L/∂θ = ∏ (local derivatives along the path)The chain rule with the forward intermediates kept. A zero anywhere in the product is a zero at the end.
Scaled dot-product attention
Attention(Q,K,V) = softmax(QKᵀ / √dₖ) VA weighted average of values, with weights computed from content. The divisor is what keeps the softmax out of saturation.
Multi-head attention
MHA(X) = [head₁ … head_h] W_O, headᵢ = Attention(XWQᵢ, XWKᵢ, XWVᵢ)Heads are slices of one budget of width d, not copies of it.
A transformer block
x ← x + Attn(Norm(x)); x ← x + MLP(Norm(x))Mix across positions, then compute within one. The additions are the residual stream.
Causal mask
A = softmax(QKᵀ/√d + M), Mᵢⱼ = 0 for j ≤ i, else −∞Added before the softmax, so masked positions get exactly zero weight rather than a small one.
Rotary position
⟨R(mθ)q, R(nθ)k⟩ = ⟨q, R((n−m)θ)k⟩The absolute indices cancel. Only the distance survives the inner product.
The vocabulary
- Parameter
- A number the model carries and training may change. Everything else about the model is fixed before the first example arrives.
- Tensor
- An array whose axes carry the meaning. The same numbers along different axes are different data.
- Logit
- An unnormalised score, before a softmax or a sigmoid. Logits have no scale of their own, which is why temperature exists.
- Token
- One integer produced by the tokeniser. The model never sees text — only a sequence of these.
- Embedding
- A row of a learned table, retrieved by token id. Retrieval is indexing; the geometry between rows is a consequence of training.
- Residual stream
- The running vector each block reads from and adds back into. A single shared communication channel of width d.
- Head
- One slice of the attention budget, of width d/h, with its own query, key and value projections.
- Receptive field
- The set of input positions that can influence one output. Fixed for a convolution, the whole sequence for attention.
- Path length
- How many operations separate two positions in the computational graph. The quantity that decides whether a long-range dependency is learnable.
- Inductive bias
- What the architecture makes easy, hard or impossible before any training happens.
- Generalisation gap
- The distance between training and held-out performance. The part of the fit that was memorisation.
- Leakage
- Any route by which held-out data influenced the model. It always moves the number in the flattering direction.
Where it is usually got wrong
said Deeper is more expressive.
in fact Only if there is a non-linearity between the layers. Stacked affine maps collapse into one affine map exactly, whatever the depth.
said Attention means the model focuses on the important tokens.
in fact Attention computes a weighted average whose weights come from content similarity. Whether the result is important is a separate question, and attention weights are weak evidence about what the model used.
said A transformer understands word order.
in fact Attention is permutation-equivariant and has no opinion about order at all. Any order-sensitivity is supplied by the positional encoding, which is where to look when it fails.
said A transformer is mostly attention.
in fact By parameter count the feed-forward blocks are usually about twice the size of the attention blocks. Attention is where the mixing happens, not where most of the model is.
said The √dₖ in attention is a cosmetic normalisation.
in fact Without it the logits grow with head width, the softmax saturates and the gradient to the query and key projections vanishes. The model still trains — just badly, and silently.
said Low training loss means the model learned the task.
in fact A large enough network reaches zero training loss on random labels. Training loss measures capacity, not understanding.
said Adam removes the need to tune the learning rate.
in fact Adam widens the usable band. Any result that changes character when the rate moves by a factor of two was never about the architecture.
said A model with a 128k context window can use 128k tokens.
in fact The window is a buffer size. What matters is what fraction of training contained sequences of that length, and it is testable in ten minutes by putting the answer at the far end.
said Mixed precision is a speed optimisation.
in fact It is first a range problem. Half precision has ample mantissa and a narrow exponent, and late-training gradients fall through its floor to zero.
The problem set
Cross-chapter problems, which is where the mathematics of a book stops being sectional. Not yet written for this book.
Book II problem set →Laboratory
Things to go and run. A worked problem is checked against arithmetic; these are checked against a machine.
Tokenise the same paragraph with two different tokenisers and count the tokens. Then find a string the two segment differently and explain what the model can no longer distinguish under each.
Chapter III. Numbers and code are the fastest places to find one.
Implement scaled dot-product attention from the equation alone, in about fifteen lines, and check it against a library implementation on random inputs.
Chapter V. Then remove the √dₖ and watch the attention distribution collapse as you raise the head width.
Take a small trained transformer, shuffle the input tokens and compare the outputs with and without the positional encoding.
Chapter VI. Without it, the outputs should be a permutation of one another.
Count the parameters of one transformer block by hand for a given d and h, split between attention and the feed-forward network. Compare with what the framework reports.
Chapter VII. If the two disagree, the disagreement is the exercise.
Plot the attention score matrix memory footprint against sequence length for a real configuration, and find the length at which it exceeds the model weights.
Chapter V, Proposition 5. Do the same for the KV cache at inference.
Take any dataset with a natural grouping — patient, site, author, session — and evaluate the same model under a random split and a grouped split.
Chapter VIII, Proposition 5. The difference between the two numbers is the size of the problem.
Write down the objective you actually care about for a project of your own, then write the loss you would train with, and list what the second omits.
Chapter I, Proposition 4. The list is where the failures will come from.
Read these
- Attention Is All You NeedThe architecture, and a trade-off table that is more instructive than the results.
- Neural Machine Translation by Jointly Learning to Align and TranslateAttention before transformers, introduced to fix one specific bottleneck. Read it first.
- Deep Residual Learning for Image RecognitionWhere the residual connection comes from, and why depth became trainable.
- Neural Machine Translation of Rare Words with Subword UnitsByte-pair encoding. Four lines of algorithm that decide what every model downstream can distinguish.
- On Layer Normalization in the Transformer ArchitectureWhy pre-norm trains without a warm-up, worked through rather than asserted.
- RoFormer: Enhanced Transformer with Rotary Position EmbeddingPosition as rotation, and the cancellation that makes it relative.
- Understanding Deep Learning Requires Rethinking GeneralizationNetworks fit random labels. Everything you believed about capacity and generalisation has to survive that.
- A Mathematical Framework for Transformer CircuitsThe residual stream taken seriously as the object of study.
- FlashAttentionThe same attention, computed with fewer memory reads. A lesson in where the cost actually was.
What comes next
Everything after this book is an answer to the question of what the sequence is made of. Book II replaces tokens with image patches and asks what convolution was giving you that attention has to earn back. Book III keeps the tokens and asks what happens when the corpus is the internet and the objective is the next word. Book IV runs both at once and has to decide how two representations are made comparable. The block does not change. What changes is what is fed to it, and what it is being asked to predict.