Md. Asif Uddin

    Chapter 5 II.5

    Position and Order

    Self-attention is permutation-equivariant, so order has to be supplied separately or it is not there at all.

    Why attention loses order, and every repair that has been tried.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics2/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Attention computes weights from content alone, which means it cannot tell a sentence from a shuffled copy of it. Order is information the mechanism structurally discards, and putting it back is a design decision with consequences for extrapolation.

    Sinusoidal position encoding as a bank of clocksThree sine waves of geometrically increasing wavelength plotted against position. Together they give every position a distinct signature: the fast wave distinguishes adjacent positions, the slow one distinguishes distant regions of the sequence.one dimension per wavelengthfast — separates neighboursmiddleslow — separates regionsone position, read off every clock at once
    Fig. 5 — Sinusoidal encoding as a bank of clocks. Fast dimensions separate neighbouring positions, slow ones separate distant regions, and together they give each position a signature.

    What this chapter covers

    • Why attention loses order
    • Positional encodings
    • Sinusoidal encoding
    • Learned positions
    • Relative position
    • Rotary position embeddings
    • Position extrapolation

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    The identity, the inverse and the transpose 0.LA.02 · Inner products, norms and cosine similarity 0.LA.03 · Projections and orthogonality 0.LA.05

    Notation

    • TSequence length in tokens
    • dModel width
    • QQuery matrix, T×d_k
    • KKey matrix, T×d_k
    • qA single query vector
    • kA single key vector

    Propositions

    1. Prop. 1Attention is permutation-invariant, and positional encoding is the repair.Self-attention is permutation-equivariant: reorder the input and the output is merely reordered. Positional encoding is what makes order visible to the model.
    2. Prop. 2A sinusoidal encoding is a bank of clocks at geometrically spaced rates.Each dimension is a sine of position at a different wavelength. Fast dimensions separate neighbours, slow ones separate regions, and together they give every position a distinct signature.
    3. Prop. 3Rotary encoding turns position into an angle, so the absolute indices cancel.Rotating the query and the key by angles proportional to their positions leaves an inner product that depends only on the difference between them. Relative position falls out of the algebra rather than being added on.
    4. Prop. 4Beyond the training length there is no row, and no experience either.A learned position table simply ends. A periodic scheme continues, but into a region the model was never trained on. Long context is a claim about training, not about the encoding.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.