Chapter 8 II.8
Training at Scale
Training a large model is an accounting problem before it is a modelling problem, and every term in the account is computable in advance.
At scale the binding constraint stops being mathematics and becomes memory and bandwidth.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
The optimisers of I.7 assume the model fits in memory and the gradient is exact. At a billion parameters neither holds, and the fixes — precision, checkpointing, clipping, parallelism — each change the arithmetic in a way worth deriving rather than trusting.
What this chapter covers
- Dataset
- Batch
- Epoch
- Optimisation
- Learning-rate schedules
- Weight decay
- Gradient clipping
- Mixed precision
- Checkpoints
- Evaluation
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Floating point 0.NU.01 · Log-sum-exp 0.NU.02 · Estimators, bias and variance 0.ST.01 · Variance and covariance 0.PR.03
Notation
- NParameter count
- BBatch size
- TSequence length in tokens
- dModel width
- LNumber of layers
- ηLearning rate
- ∇Gradient operator
- VarVariance
Propositions
- Prop. 1Mixed precision is a question of range before it is a question of speed.Half precision has ample precision and a narrow exponent. Small gradients fall through its floor and round to zero, which is what loss scaling exists to prevent.
- Prop. 2Weight decay and gradient clipping guard different failures and do not substitute for each other.Clipping bounds the size of a step and does nothing on almost every one. Decay pulls every parameter toward zero on every step. One is about the update, the other about where the parameters end up.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.