Book II
Foundations
tokenisation, attention, the transformer block
Build the conceptual and mathematical footing the rest of the corpus stands on. By the end you should understand how a network represents information and how a transformer processes a sequence.
8 chapters22 propositions written
Read first
Mathematics assumed — follow these when a step stops making sense.
Vectors, matrices and the row-major convention 0.LA.01 · The identity, the inverse and the transpose 0.LA.02 · The gradient, and the layout convention 0.MC.01 · The chain rule 0.MC.03 · Expectation 0.PR.02 · Entropy 0.IT.01 · Convexity 0.OP.01
- Chapter 1II.1Tokenisation and EmbeddingsTokenisation fixes, before any learning happens, which distinctions the model is able to make.Text as discrete symbols · Characters, words and subwords · BPE · Unigram tokenisation · Vocabulary · Token IDs · Embeddings · Embedding spaces · Contextual representations
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 2II.2From Recurrence to AttentionAttention was the answer to a specific bottleneck, not a general improvement.
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 3II.3AttentionAttention is a content-dependent weighted aggregation: each output row is a convex combination of value rows.Queries · Keys · Values · Similarity · Attention scores · Softmax · Weighted aggregation · Scaled dot-product attention · Self-attention · Cross-attention
1/5 problems1/4 variants1/10 exercisesowes 13 more
- Chapter 4II.4Multi-Head AttentionHeads are slices of one budget, not copies of one mechanism, and partitioning preserves the total cost.
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 5II.5Position and OrderSelf-attention is permutation-equivariant, so order has to be supplied separately or it is not there at all.Why attention loses order · Positional encodings · Sinusoidal encoding · Learned positions · Relative position · Rotary position embeddings · Position extrapolation
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 6II.6The Transformer BlockA transformer block mixes across positions and then computes within them, and the residual stream is what every layer edits rather than replaces.Transformer block · Multi-head attention · Feed-forward network · Residual connections · Layer normalisation · Encoder · Decoder · Encoder-decoder models · Causal masking · Full forward pass
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 7II.7Encoder, Decoder, and Encoder–DecoderEncoder, decoder and encoder–decoder differ only in what each stack may look at, and a mask is how that is enforced.
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 8II.8Training at ScaleTraining a large model is an accounting problem before it is a modelling problem, and every term in the account is computable in advance.Dataset · Batch · Epoch · Optimisation · Learning-rate schedules · Weight decay · Gradient clipping · Mixed precision · Checkpoints · Evaluation
0/5 problems0/4 variants0/10 exercisesowes 15 more
Practical connection
Tokeniser, then attention, then a full Transformer block, then a training loop that survives its own memory budget.
Your implementation must reproduce the hand-computed numbers to four decimal places before you move on. A reader who skipped the derivations cannot pass their own unit test.
Verified againstII.1.B01 · II.3.B01 · II.6.B02 · II.8.B01