Chapter 3 II.3
Attention
Attention is a content-dependent weighted aggregation: each output row is a convex combination of value rows.
Content-addressed aggregation, in full.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
The bottleneck of II.2 is removed by letting every position read every other position directly. The question this chapter answers is how a position decides what to read, using only content and never an index.
What this chapter covers
- Queries
- Keys
- Values
- Similarity
- Attention scores
- Softmax
- Weighted aggregation
- Scaled dot-product attention
- Self-attention
- Cross-attention
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Inner products, norms and cosine similarity 0.LA.03 · The softmax Jacobian 0.MC.06 · Variance and covariance 0.PR.03 · Log-sum-exp 0.NU.02
Notation
- QQuery matrix, T×d_k
- KKey matrix, T×d_k
- VValue matrix, T×d_v
- d_kKey and query width inside one attention head
- d_vValue width inside one attention head
- TSequence length in tokens
- qA single query vector
- kA single key vector
- τTemperature, in a softmax or a contrastive loss
- softmaxThe normalised exponential, applied row-wise unless stated
- VarVariance
- ∇Gradient operator
Propositions
- Prop. 1Attention is a weighted average whose weights are computed from content.Attention outputs a convex combination of value vectors, weighted by how well each key matches the query. The weights come from content; the position of a token plays no part.
- Prop. 2Query, key and value are three different questions put to the same vector.One token vector is projected three ways. The query is what it is looking for, the key is what it advertises, the value is what it hands over once matched — and only the value reaches the output.
- Prop. 3The divisor in scaled dot-product attention is what keeps the gradient alive.Dot products of d-dimensional vectors have variance proportional to d. Left unscaled the logits grow with head width, the softmax saturates, and the gradient through it goes to zero.
- Prop. 4Self-attention and cross-attention are one operation under two wirings.The block is identical. Only the source of the keys and values changes — the same sequence as the queries, or a different one.
- Prop. 5Constant path length is bought with a quadratic score matrix.Every pair of positions gets a score, so time and memory grow as the square of the sequence length. This is the price of the property that makes attention worth having.
Worked problems
1/5 problems1/4 variants1/10 exercisesowes 13 more