Md. Asif Uddin

    Chapter 6 VIII.6

    Training Systems

    A training run is bounded by memory and bandwidth before it is bounded by ideas, and both bounds are measurable rather than guessed.

    Measurement and diagnosis: where the time actually goes.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics2/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    The budget was derived in II.8. What remains is the gap between that budget and what the hardware actually achieves, which is a measurement problem with a small number of standard causes.

    Fewer operations, more time, same accuracyTwo models that both reach 84.0% on ImageNet. One uses 1.8 times fewer floating-point operations and runs 2.7 times slower on the same hardware, because depthwise convolutions are bound by memory bandwidth rather than arithmetic.both land at 84.0% on ImageNetEfficientNet-B6FLOPsarithmetictimewall clockResNet-RS-350FLOPsarithmetictimewall clockAccelerators are bound by memory bandwidth, not arithmetic. "Efficient" meant FLOP-efficient, and a decadeof practitioners read it as fast.
    Fig. 6 — Two models at the same accuracy: one with 1.8x fewer operations and 2.7x more wall-clock time. Accelerators are bound by memory bandwidth, not arithmetic.

    What this chapter covers

    • GPU
    • Memory
    • Batch size
    • Mixed precision
    • Distributed training
    • Checkpointing
    • Profiling
    • Inference

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Floating point 0.NU.01 · Log-sum-exp 0.NU.02 · Vectors, matrices and the row-major convention 0.LA.01

    Notation

    • NParameter count
    • BBatch size
    • TSequence length in tokens
    • dModel width
    • LNumber of layers
    • ηLearning rate

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.