Md. Asif Uddin

    Chapter 3 II.3

    Attention

    Attention is a content-dependent weighted aggregation: each output row is a convex combination of value rows.

    Content-addressed aggregation, in full.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics2/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    The bottleneck of II.2 is removed by letting every position read every other position directly. The question this chapter answers is how a position decides what to read, using only content and never an index.

    Attention as a weighted average over valuesA single query is compared against four keys. The resulting softmax weights — 0.06, 0.61, 0.09 and 0.24 — are shown as horizontal bars, and the output is the sum of the value vectors scaled by those weights. The weights come from content, not from position.query qsoftmax(q · kᵢ / √d)k₁0.06k₂0.61k₃0.09k₄0.24Σ wᵢ vᵢweights are content-addressed — nothing here depends on i
    Fig. 3 — Attention as a weighted average over values, with weights computed by comparing a query against every key. Nothing in the computation depends on position.

    What this chapter covers

    • Queries
    • Keys
    • Values
    • Similarity
    • Attention scores
    • Softmax
    • Weighted aggregation
    • Scaled dot-product attention
    • Self-attention
    • Cross-attention

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Inner products, norms and cosine similarity 0.LA.03 · The softmax Jacobian 0.MC.06 · Variance and covariance 0.PR.03 · Log-sum-exp 0.NU.02

    Notation

    • QQuery matrix, T×d_k
    • KKey matrix, T×d_k
    • VValue matrix, T×d_v
    • d_kKey and query width inside one attention head
    • d_vValue width inside one attention head
    • TSequence length in tokens
    • qA single query vector
    • kA single key vector
    • τTemperature, in a softmax or a contrastive loss
    • softmaxThe normalised exponential, applied row-wise unless stated
    • VarVariance
    • ∇Gradient operator

    Propositions

    1. Prop. 1Attention is a weighted average whose weights are computed from content.Attention outputs a convex combination of value vectors, weighted by how well each key matches the query. The weights come from content; the position of a token plays no part.
    2. Prop. 2Query, key and value are three different questions put to the same vector.One token vector is projected three ways. The query is what it is looking for, the key is what it advertises, the value is what it hands over once matched — and only the value reaches the output.
    3. Prop. 3The divisor in scaled dot-product attention is what keeps the gradient alive.Dot products of d-dimensional vectors have variance proportional to d. Left unscaled the logits grow with head width, the softmax saturates, and the gradient through it goes to zero.
    4. Prop. 4Self-attention and cross-attention are one operation under two wirings.The block is identical. Only the source of the keys and values changes — the same sequence as the queries, or a different one.
    5. Prop. 5Constant path length is bought with a quadratic score matrix.Every pair of positions gets a score, so time and memory grow as the square of the sequence length. This is the price of the property that makes attention worth having.

    Worked problems

    1/5 problems1/4 variants1/10 exercisesowes 13 more

    1. II.3.B01 — Attention on three tokens, by handnumeric▲▲△

    The whole problem set, with the exercises →