Md. Asif Uddin

    Chapter 8 II.8

    Training at Scale

    Training a large model is an accounting problem before it is a modelling problem, and every term in the account is computable in advance.

    At scale the binding constraint stops being mathematics and becomes memory and bandwidth.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    The optimisers of I.7 assume the model fits in memory and the gradient is exact. At a billion parameters neither holds, and the fixes — precision, checkpointing, clipping, parallelism — each change the arithmetic in a way worth deriving rather than trusting.

    Weight decay and gradient clipping guard different thingsA bar chart of gradient norms over fourteen steps, mostly small with two enormous spikes. A horizontal clipping threshold cuts only the two spikes. Beneath, a separate note shows weight decay acting on every step regardless.gradient norm per stepclip thresholdClipping is a rare intervention: on almost every step it does nothing at all.Weight decay is the opposite: a small constant pull toward zero, applied on every step to everyparameter, whatever the gradient is doing. One bounds the step, the other bounds where they land.
    Fig. 8 — Gradient norms over fourteen steps with a clipping threshold. Clipping touches only the rare spike; weight decay pulls on every parameter every step. They guard different failures.

    What this chapter covers

    • Dataset
    • Batch
    • Epoch
    • Optimisation
    • Learning-rate schedules
    • Weight decay
    • Gradient clipping
    • Mixed precision
    • Checkpoints
    • Evaluation

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Floating point 0.NU.01 · Log-sum-exp 0.NU.02 · Estimators, bias and variance 0.ST.01 · Variance and covariance 0.PR.03

    Notation

    • NParameter count
    • BBatch size
    • TSequence length in tokens
    • dModel width
    • LNumber of layers
    • ηLearning rate
    • ∇Gradient operator
    • VarVariance

    Propositions

    1. Prop. 1Mixed precision is a question of range before it is a question of speed.Half precision has ample precision and a narrow exponent. Small gradients fall through its floor and round to zero, which is what loss scaling exists to prevent.
    2. Prop. 2Weight decay and gradient clipping guard different failures and do not substitute for each other.Clipping bounds the size of a step and does nothing on almost every one. Decay pulls every parameter toward zero on every step. One is about the update, the other about where the parameters end up.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.