Md. Asif Uddin

    Chapter 4 IV.4

    Scaling

    Loss falls as a power law in parameters and data, and a fixed compute budget has one optimal split between them.

    Loss falls as a power law, and compute decides how to split it.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Given more compute, should the model be larger or the corpus? The question has an answer with a derivation behind it, and the derivation needs the cost model of II.8 rather than a fresh one.

    Zero-shot average across 27 benchmarksThree bars on a truncated axis. Going from eight billion parameters to eighteen billion gains seven tenths of a point. Raising the input resolution on the smaller model, with no extra parameters at all, gains six.zero-shot average, 27 benchmarks79.4%EVA-CLIP-8B224px80%EVA-CLIP-8B448px, same params80.7%EVA-CLIP-18B224px, 2.2x params78.5Axis truncated. The data stayed fixed the whole way — scaling parameters still works, it just works slowly,and the lever everyone else is pulling is the data.
    Fig. 4 — Doubling the parameters on a fixed dataset bought seven tenths of a point. Raising the input resolution, with no extra parameters at all, bought six.

    What this chapter covers

    • Parameter count
    • Data
    • Compute
    • Scaling laws
    • Compute and data trade-offs
    • Inference scaling

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Lagrange multipliers 0.OP.04 · Estimators, bias and variance 0.ST.01

    Notation

    • NParameter count
    • DDataset size in tokens
    • CChannels in vision; compute in FLOPs in the scaling chapters
    • ℒThe loss
    • 𝔼Expectation

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.