Md. Asif Uddin

    Book II

    Foundations

    tokenisation, attention, the transformer block

    Build the conceptual and mathematical footing the rest of the corpus stands on. By the end you should understand how a network represents information and how a transformer processes a sequence.

    8 chapters22 propositions written

    Read first

    neural networks

    Mathematics assumed — follow these when a step stops making sense.

    Vectors, matrices and the row-major convention 0.LA.01 · The identity, the inverse and the transpose 0.LA.02 · The gradient, and the layout convention 0.MC.01 · The chain rule 0.MC.03 · Expectation 0.PR.02 · Entropy 0.IT.01 · Convexity 0.OP.01

    1. Chapter 1II.1Tokenisation and EmbeddingsTokenisation fixes, before any learning happens, which distinctions the model is able to make.Text as discrete symbols · Characters, words and subwords · BPE · Unigram tokenisation · Vocabulary · Token IDs · Embeddings · Embedding spaces · Contextual representationsM24 written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    2. Chapter 2II.2From Recurrence to AttentionAttention was the answer to a specific bottleneck, not a general improvement.M21 written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    3. Chapter 3II.3AttentionAttention is a content-dependent weighted aggregation: each output row is a convex combination of value rows.Queries · Keys · Values · Similarity · Attention scores · Softmax · Weighted aggregation · Scaled dot-product attention · Self-attention · Cross-attentionM35 written

      1/5 problems1/4 variants1/10 exercisesowes 13 more

    4. Chapter 4II.4Multi-Head AttentionHeads are slices of one budget, not copies of one mechanism, and partitioning preserves the total cost.M31 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    5. Chapter 5II.5Position and OrderSelf-attention is permutation-equivariant, so order has to be supplied separately or it is not there at all.Why attention loses order · Positional encodings · Sinusoidal encoding · Learned positions · Relative position · Rotary position embeddings · Position extrapolationM34 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    6. Chapter 6II.6The Transformer BlockA transformer block mixes across positions and then computes within them, and the residual stream is what every layer edits rather than replaces.Transformer block · Multi-head attention · Feed-forward network · Residual connections · Layer normalisation · Encoder · Decoder · Encoder-decoder models · Causal masking · Full forward passM33 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    7. Chapter 7II.7Encoder, Decoder, and Encoder–DecoderEncoder, decoder and encoder–decoder differ only in what each stack may look at, and a mask is how that is enforced.M32 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    8. Chapter 8II.8Training at ScaleTraining a large model is an accounting problem before it is a modelling problem, and every term in the account is computable in advance.Dataset · Batch · Epoch · Optimisation · Learning-rate schedules · Weight decay · Gradient clipping · Mixed precision · Checkpoints · EvaluationM32 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    Problem setWhere the chapters have to be used togetherNot yet written. A book with load-bearing chapters owes at least eight cross-chapter problems and two that reach back into an earlier book.

    Practical connection

    Tokeniser, then attention, then a full Transformer block, then a training loop that survives its own memory budget.

    Your implementation must reproduce the hand-computed numbers to four decimal places before you move on. A reader who skipped the derivations cannot pass their own unit test.

    Verified againstII.1.B01 · II.3.B01 · II.6.B02 · II.8.B01

    ClosingThe whole book on one plateThe map, the key equations with their glosses, what you should now be able to derive unaided, the vocabulary, the misconceptions and the laboratory. Read it after the chapters, then again before the next book.