Chapter 4 II.4
Multi-Head Attention
Heads are slices of one budget, not copies of one mechanism, and partitioning preserves the total cost.
One attention pattern per relation type; several relations are needed at once.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
A single attention distribution must commit to one notion of relevance per position. Syntax and coreference are different relations, and a model that can only express one of them at a time cannot represent both.
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Vectors, matrices and the row-major convention 0.LA.01 · The chain rule for matrix products 0.MC.07
Notation
- hNumber of attention heads
- dModel width
- d_kKey and query width inside one attention head
- QQuery matrix, T×d_k
- KKey matrix, T×d_k
- VValue matrix, T×d_v
- TSequence length in tokens
Propositions
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.