Attention on three tokens, by hand
numeric▲▲△All values rounded to 4 d.p. Intermediate quantities carry full precision; only what is printed is rounded.
STATEMENT
Three tokens attend to one another under scaled dot-product attention. Compute the whole forward pass by hand and confirm that every attention row is a probability distribution.
GIVEN
Row of each matrix belongs to token . Row-major throughout, as fixed in the Apparatus.
FIND
The score matrix , the scaled scores , the attention weights , and the output — each a matrix except , which is .
STRATEGY
Name each intermediate rather than nesting them, so a wrong digit can be traced to the step that produced it.
SOLUTION
Step 1 — the scores. , so is the inner product of query row with key row .
Token 3’s query is , which agrees with both basis directions, so its row is the largest. Nothing about position enters: is a function of content alone.
Step 2 — the scaling. Divide by .
Step 3 — the softmax, row by row. Take row 1, . Subtract the row maximum first, which changes nothing and removes the overflow (Log-sum-exp 0.NU.02):
The denominator is , so
Row 2 is row 1 with the first two entries exchanged, by the symmetry of the inputs. Row 3 has ; subtracting the maximum gives , exponentials , denominator , and
Step 4 — the aggregation. . For row 1, column 1:
and column 2 is .
Answer
is and dimensionless; is and carries the units of .
Check — numeric · ii-3-b01-attention-by-hand.py
S = [[sum(q[i] * k[i] for i in range(d_k)) for k in K] for q in Q]
Z = [[s / sqrt(d_k) for s in row] for row in S]
A = [softmax(row) for row in Z]
O = [[sum(a[j] * V[j][c] for j in range(3)) for c in range(2)] for a in A]The full twenty-line script is scripts/snippets/ii-3-b01-attention-by-hand.py,
and its output is diffed against the digits above on every build.
Executed in CI. The digits above are the digits it printed.
Check — sanity
Three independent reasons the answer is right.
Rows sum to one. , and likewise for the other two. A softmax row that does not sum to one is an arithmetic slip, not a modelling choice.
Every output lies inside the convex hull of . The value rows are , and ; every entry of lies in , as a convex combination must. This is the whole content of Proposition 1: attention mixes, and cannot leave the hull.
The symmetry survives. here, so is symmetric, and rows 1 and 2 of the answer are exchanged copies of one another. They are.
Where this breaks
The convex-hull argument holds only because the weights are non-negative and sum to one. Remove the softmax — score-weighted sums, or normalisation by anything other than the total — and the output can leave the hull entirely. That is why the feed-forward block after attention is not decoration.
Variation
Recompute with and predict before doing any arithmetic. Then set with the same and and say what happens to the sharpness of , and why.