Proposition 1II.3.P0143 of 86 in the corpus
Attention is a weighted average whose weights are computed from content.
Attention outputs a convex combination of value vectors, weighted by how well each key matches the query. The weights come from content; the position of a token plays no part.
Demonstration
Attention is often introduced as a mechanism for “focus”, which is evocative and not very useful. Mechanically it is an average.
For a query vector q and a set of key–value pairs (kᵢ, vᵢ):
wᵢ = softmax( q · kᵢ / √d )
out = Σ wᵢ vᵢ
The output is a convex combination of the value vectors. It lives in the convex hull of the values and can never leave it — which is why the feed-forward block after attention is not optional decoration. Attention mixes; the MLP transforms.
The division by √d is not cosmetic. For q and k with unit-variance components, the dot product has variance d. Left unscaled, the logits grow with dimension, the softmax saturates, and gradients through it vanish. Dividing by √d restores unit variance and keeps the distribution soft enough to train.
The decisive property is where the weights come from: content. The index i appears nowhere in the computation of wᵢ. Two identical keys receive identical weights no matter where in the sequence they sit.
Corollary
Because the weights are content-addressed and index-free, the whole operation is blind to order. That blindness is the subject of Chapter VI, and the whole of that chapter is the repair.
Sources
Depends on
Used by
- II.3.P02 — Query, key and value are three different questions put to the same vector.
- II.3.P03 — The divisor in scaled dot-product attention is what keeps the gradient alive.
- II.3.P05 — Constant path length is bought with a quadratic score matrix.
- II.5.P01 — Attention is permutation-invariant, and positional encoding is the repair.
- II.6.P01 — A transformer block mixes across positions, then computes within one.
- II.7.P01 — A causal mask is what turns a transformer into a language model.