Chapter 6 II.6
The Transformer Block
A transformer block mixes across positions and then computes within them, and the residual stream is what every layer edits rather than replaces.
Two sublayers, each adding to a running representation rather than replacing it.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Attention mixes but cannot transform: its output is trapped in the convex hull of its values. A block is the smallest unit that both mixes and computes, and the residual connection is what lets many of them stack.
What this chapter covers
- Transformer block
- Multi-head attention
- Feed-forward network
- Residual connections
- Layer normalisation
- Encoder
- Decoder
- Encoder-decoder models
- Causal masking
- Full forward pass
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
The identity, the inverse and the transpose 0.LA.02 · The chain rule for matrix products 0.MC.07 · Log-sum-exp 0.NU.02
Notation
- dModel width
- TSequence length in tokens
- LNumber of layers
- hNumber of attention heads
- NParameter count
- WA weight matrix
- IThe identity matrix
- JA Jacobian matrix
- BBatch size
Propositions
- Prop. 1A transformer block mixes across positions, then computes within one.Attention is the only part of a block that moves information between tokens. The feed-forward network sees each position alone, and the two alternate for the whole depth of the model.
- Prop. 2Layers write into a shared stream rather than replacing it.Every sublayer adds its output to the running representation. This is why gradients reach the first layer, and why a state from partway up the stack can be read with the output head.
- Prop. 3Where the normalisation sits decides whether the shortcut is clean.Post-norm places a normalisation on the residual path and needs a warm-up to train deeply. Pre-norm leaves the path untouched and trains from a flat start. Same components, different order.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.