Chapter 2 IV.2
Autoregressive Transformers
Prefill is compute-bound and decode is memory-bandwidth-bound, and the cache is what separates them.
Generation is a different computation from training.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Training runs the whole sequence at once; generation runs one token at a time. The arithmetic of those two regimes differs by orders of magnitude in intensity, and almost every inference optimisation is a response to that gap.
What this chapter covers
- Causal masking
- Decoder-only architecture
- GPT-style models
- Context windows
- KV cache
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Log-sum-exp 0.NU.02 · Distributions, discrete and continuous 0.PR.01
Notation
- TSequence length in tokens
- dModel width
- LNumber of layers
- hNumber of attention heads
- BBatch size
- τTemperature, in a softmax or a contrastive loss
- softmaxThe normalised exponential, applied row-wise unless stated
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.