Md. Asif Uddin

    Book II · Closing

    Foundations, gathered

    8 chapters · 22 propositions

    The whole of Book I is one plate. Text becomes tokens, tokens become vectors, position is added or applied, and then the same block runs some number of times: mix across positions, compute within one, add the result back into the stream. At the end the stream is projected back to the vocabulary and a distribution comes out.

    If any stage of that map is one you could not explain to somebody else, the chapter is named beside it.

    The propositions were written to be read in order, because they depend on one another in a specific way: a tensor is only interesting once you know a model is a function of parameters; attention only makes sense once you know why a recurrence could not do it; the residual stream only matters once you have seen a long product of Jacobians decay. The dependency graph on each page is the real structure, and it is worth walking backwards from any proposition that does not land.

    One thing to carry forward before the next book. Of everything here, the proposition that constrains all the others is the last one. Architecture, loss, schedule and precision are judged by a number, and the number is worth exactly what the split makes it worth.

    The whole of Book I on one plateA vertical map from raw text to output logits. Text is tokenised, embedded and given position, then passed through a stack of transformer blocks, each containing normalisation, attention, a residual addition, a second normalisation and a feed-forward network. The final state is unembedded to logits. Each stage is labelled with the chapter that covers it.from characters to logitstexttokeniserCh. IIItoken idsCh. IIIembedding + positionCh. III, VI× N blocksnormCh. VIImulti-head attentionCh. V, VIIresidual addCh. VIInormCh. VIIfeed-forwardCh. II, VIIresidual addCh. VIIstreamunembed → logitsCh. I, VIIIEverything in Book I ison this plate. If a stagehere is not one you couldexplain to somebody else,that is the chapter toread again.
    Fig. — — The whole of Book I on one plate, from characters to logits, with the chapter covering each stage named beside it.

    The equations

    1. A model

      ŷ = f(x ; θ)

      Two arguments. The world supplies the first; training may only change the second.

    2. A linear layer

      y = Wx + b

      Affine, and affine maps compose into a single affine map. Depth without a non-linearity buys nothing.

    3. The gradient step

      θ ← θ − η ∇θ L

      The gradient gives the direction; η is where the distance has to come from.

    4. Backpropagation

      ∂L/∂θ = ∏ (local derivatives along the path)

      The chain rule with the forward intermediates kept. A zero anywhere in the product is a zero at the end.

    5. Scaled dot-product attention

      Attention(Q,K,V) = softmax(QKᵀ / √dₖ) V

      A weighted average of values, with weights computed from content. The divisor is what keeps the softmax out of saturation.

    6. Multi-head attention

      MHA(X) = [head₁ … head_h] W_O,   headᵢ = Attention(XWQᵢ, XWKᵢ, XWVᵢ)

      Heads are slices of one budget of width d, not copies of it.

    7. A transformer block

      x ← x + Attn(Norm(x));  x ← x + MLP(Norm(x))

      Mix across positions, then compute within one. The additions are the residual stream.

    8. Causal mask

      A = softmax(QKᵀ/√d + M),   Mᵢⱼ = 0 for j ≤ i, else −∞

      Added before the softmax, so masked positions get exactly zero weight rather than a small one.

    9. Rotary position

      ⟨R(mθ)q, R(nθ)k⟩ = ⟨q, R((n−m)θ)k⟩

      The absolute indices cancel. Only the distance survives the inner product.

    The vocabulary

    Parameter
    A number the model carries and training may change. Everything else about the model is fixed before the first example arrives.
    Tensor
    An array whose axes carry the meaning. The same numbers along different axes are different data.
    Logit
    An unnormalised score, before a softmax or a sigmoid. Logits have no scale of their own, which is why temperature exists.
    Token
    One integer produced by the tokeniser. The model never sees text — only a sequence of these.
    Embedding
    A row of a learned table, retrieved by token id. Retrieval is indexing; the geometry between rows is a consequence of training.
    Residual stream
    The running vector each block reads from and adds back into. A single shared communication channel of width d.
    Head
    One slice of the attention budget, of width d/h, with its own query, key and value projections.
    Receptive field
    The set of input positions that can influence one output. Fixed for a convolution, the whole sequence for attention.
    Path length
    How many operations separate two positions in the computational graph. The quantity that decides whether a long-range dependency is learnable.
    Inductive bias
    What the architecture makes easy, hard or impossible before any training happens.
    Generalisation gap
    The distance between training and held-out performance. The part of the fit that was memorisation.
    Leakage
    Any route by which held-out data influenced the model. It always moves the number in the flattering direction.

    Where it is usually got wrong

    • said Deeper is more expressive.

      in fact Only if there is a non-linearity between the layers. Stacked affine maps collapse into one affine map exactly, whatever the depth.

    • said Attention means the model focuses on the important tokens.

      in fact Attention computes a weighted average whose weights come from content similarity. Whether the result is important is a separate question, and attention weights are weak evidence about what the model used.

    • said A transformer understands word order.

      in fact Attention is permutation-equivariant and has no opinion about order at all. Any order-sensitivity is supplied by the positional encoding, which is where to look when it fails.

    • said A transformer is mostly attention.

      in fact By parameter count the feed-forward blocks are usually about twice the size of the attention blocks. Attention is where the mixing happens, not where most of the model is.

    • said The √dₖ in attention is a cosmetic normalisation.

      in fact Without it the logits grow with head width, the softmax saturates and the gradient to the query and key projections vanishes. The model still trains — just badly, and silently.

    • said Low training loss means the model learned the task.

      in fact A large enough network reaches zero training loss on random labels. Training loss measures capacity, not understanding.

    • said Adam removes the need to tune the learning rate.

      in fact Adam widens the usable band. Any result that changes character when the rate moves by a factor of two was never about the architecture.

    • said A model with a 128k context window can use 128k tokens.

      in fact The window is a buffer size. What matters is what fraction of training contained sequences of that length, and it is testable in ten minutes by putting the answer at the far end.

    • said Mixed precision is a speed optimisation.

      in fact It is first a range problem. Half precision has ample mantissa and a narrow exponent, and late-training gradients fall through its floor to zero.

    The problem set

    Cross-chapter problems, which is where the mathematics of a book stops being sectional. Not yet written for this book.

    Book II problem set →

    Laboratory

    Things to go and run. A worked problem is checked against arithmetic; these are checked against a machine.

    1. Tokenise the same paragraph with two different tokenisers and count the tokens. Then find a string the two segment differently and explain what the model can no longer distinguish under each.

      Chapter III. Numbers and code are the fastest places to find one.

    2. Implement scaled dot-product attention from the equation alone, in about fifteen lines, and check it against a library implementation on random inputs.

      Chapter V. Then remove the √dₖ and watch the attention distribution collapse as you raise the head width.

    3. Take a small trained transformer, shuffle the input tokens and compare the outputs with and without the positional encoding.

      Chapter VI. Without it, the outputs should be a permutation of one another.

    4. Count the parameters of one transformer block by hand for a given d and h, split between attention and the feed-forward network. Compare with what the framework reports.

      Chapter VII. If the two disagree, the disagreement is the exercise.

    5. Plot the attention score matrix memory footprint against sequence length for a real configuration, and find the length at which it exceeds the model weights.

      Chapter V, Proposition 5. Do the same for the KV cache at inference.

    6. Take any dataset with a natural grouping — patient, site, author, session — and evaluate the same model under a random split and a grouped split.

      Chapter VIII, Proposition 5. The difference between the two numbers is the size of the problem.

    7. Write down the objective you actually care about for a project of your own, then write the loss you would train with, and list what the second omits.

      Chapter I, Proposition 4. The list is where the failures will come from.

    Read these

    What comes next

    Everything after this book is an answer to the question of what the sequence is made of. Book II replaces tokens with image patches and asks what convolution was giving you that attention has to earn back. Book III keeps the tokens and asks what happens when the corpus is the internet and the objective is the next word. Book IV runs both at once and has to decide how two representations are made comparable. The block does not change. What changes is what is fed to it, and what it is being asked to predict.