Md. Asif Uddin

    Chapter 2 IV.2

    Autoregressive Transformers

    Prefill is compute-bound and decode is memory-bandwidth-bound, and the cache is what separates them.

    Generation is a different computation from training.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Training runs the whole sequence at once; generation runs one token at a time. The arithmetic of those two regimes differs by orders of magnitude in intensity, and almost every inference optimisation is a response to that gap.

    Cost per token at a one-million-token contextTwo bars per model. Compared with the previous generation, V4 needs about twenty-seven percent of the per-token inference operations and ten percent of the key-value cache at a one-million-token context.at 1M context, relative to V3.2V3.2FLOPs100%KV cache100%V4FLOPs27%KV cache10%Compressed Sparse Attention interleaved with Heavily Compressed Attention. Everyone quoted the 1.6 trillion.The number that matters is the 10%.
    Fig. 2 — At a one-million-token context, roughly a quarter of the per-token operations and a tenth of the key-value cache. Everyone quoted the parameter count.

    What this chapter covers

    • Causal masking
    • Decoder-only architecture
    • GPT-style models
    • Context windows
    • KV cache

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Log-sum-exp 0.NU.02 · Distributions, discrete and continuous 0.PR.01

    Notation

    • TSequence length in tokens
    • dModel width
    • LNumber of layers
    • hNumber of attention heads
    • BBatch size
    • τTemperature, in a softmax or a contrastive loss
    • softmaxThe normalised exponential, applied row-wise unless stated

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.