Md. Asif Uddin

    Chapter 7 II.7

    Encoder, Decoder, and Encoder–Decoder

    Encoder, decoder and encoder–decoder differ only in what each stack may look at, and a mask is how that is enforced.

    Three architectures, one difference: who is allowed to see whom.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics2/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    The same block can be wired three ways, and the wiring decides whether the model classifies, generates, or translates. The whole difference is a matrix of zeros and minus infinities.

    The causal maskA square grid of attention scores for seven tokens. The lower triangle, including the diagonal, is open: each token may attend to itself and to everything before it. The upper triangle is closed, because those positions lie in the future.query row attends to key columnthe−∞−∞−∞−∞−∞−∞cat−∞−∞−∞−∞−∞sat−∞−∞−∞−∞on−∞−∞−∞the−∞−∞warm−∞matthecatsatonthewarmmatSet before the softmax, so themasked entries receive exactlyzero weight rather than a small one.This is what lets every position betrained at once on the same sequence:n prediction problems, one forward pass.Remove the triangle and the same weights become an encoder. Nothing else about the block changes.
    Fig. 7 — The causal mask. Setting the future to minus infinity before the softmax is what lets every position in a sequence be trained as a separate prediction in one pass.

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Distributions, discrete and continuous 0.PR.01 · Cross-entropy 0.IT.02

    Notation

    • TSequence length in tokens
    • dModel width
    • softmaxThe normalised exponential, applied row-wise unless stated
    • LNumber of layers

    Propositions

    1. Prop. 1A causal mask is what turns a transformer into a language model.Setting the future to minus infinity before the softmax lets every position in a sequence be trained as its own prediction problem, in a single forward pass.
    2. Prop. 2Encoder, decoder and encoder–decoder differ only in what each stack may look at.The block is unchanged in all three. What differs is the mask, and whether a second stack supplies the keys and values.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.