Md. Asif Uddin

    Chapter 6 II.6

    The Transformer Block

    A transformer block mixes across positions and then computes within them, and the residual stream is what every layer edits rather than replaces.

    Two sublayers, each adding to a running representation rather than replacing it.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Attention mixes but cannot transform: its output is trapped in the convex hull of its values. A block is the smallest unit that both mixes and computes, and the residual connection is what lets many of them stack.

    One transformer blockA block with two sublayers. The first normalises, applies multi-head attention across positions and adds the result back to the input. The second normalises, applies a position-wise feed-forward network and adds again.mix across positionsthen compute within onexoutnormmulti-head attention+the residual path — the block edits, it does not replacenormfeed-forward (per token)+no token sees another here
    Fig. 6 — One transformer block: attention moves information between positions, the feed-forward network does the work within one, and both are wrapped in a residual addition.

    What this chapter covers

    • Transformer block
    • Multi-head attention
    • Feed-forward network
    • Residual connections
    • Layer normalisation
    • Encoder
    • Decoder
    • Encoder-decoder models
    • Causal masking
    • Full forward pass

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    The identity, the inverse and the transpose 0.LA.02 · The chain rule for matrix products 0.MC.07 · Log-sum-exp 0.NU.02

    Notation

    • dModel width
    • TSequence length in tokens
    • LNumber of layers
    • hNumber of attention heads
    • NParameter count
    • WA weight matrix
    • IThe identity matrix
    • JA Jacobian matrix
    • BBatch size

    Propositions

    1. Prop. 1A transformer block mixes across positions, then computes within one.Attention is the only part of a block that moves information between tokens. The feed-forward network sees each position alone, and the two alternate for the whole depth of the model.
    2. Prop. 2Layers write into a shared stream rather than replacing it.Every sublayer adds its output to the running representation. This is why gradients reach the first layer, and why a state from partway up the stack can be read with the output head.
    3. Prop. 3Where the normalisation sits decides whether the shortcut is clean.Post-norm places a normalisation on the residual path and needs a warm-up to train deeply. Pre-norm leaves the path untouched and trains from a flat start. Same components, different order.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.