Md. Asif Uddin

    Chapter 4 II.4

    Multi-Head Attention

    Heads are slices of one budget, not copies of one mechanism, and partitioning preserves the total cost.

    One attention pattern per relation type; several relations are needed at once.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics2/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    A single attention distribution must commit to one notion of relevance per position. Syntax and coreference are different relations, and a model that can only express one of them at a time cannot represent both.

    One attention budget, divided into four headsA vector of width d split into four slices of width d over four. Each slice runs its own attention and attends to a different relation. The four outputs are concatenated back to width d and mixed by an output projection.one vector, width dhead 1d / 4previous tokenhead 2d / 4the subjecthead 3d / 4matching quotehead 4d / 4nothing usefulconcatenate, then W_OHeads are cheap because they are slices. What they buy is several relations at once;what they cost is resolution inside each.
    Fig. 4 — One attention budget divided into four heads. Heads are slices rather than copies, so several relations are attended to at once at the cost of resolution inside each.

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Vectors, matrices and the row-major convention 0.LA.01 · The chain rule for matrix products 0.MC.07

    Notation

    • hNumber of attention heads
    • dModel width
    • d_kKey and query width inside one attention head
    • QQuery matrix, T×d_k
    • KKey matrix, T×d_k
    • VValue matrix, T×d_v
    • TSequence length in tokens

    Propositions

    1. Prop. 1Heads are slices of one budget, not copies of one mechanism.Multi-head attention splits the width into parts and runs attention in each. It costs what single-head attention costs, and buys several relations at once at the price of resolution inside each.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.