Md. Asif Uddin

    Chapter 3 IV.3

    Pretraining

    A pretraining corpus is a set of decisions about duplication, mixture and budget, and each one is measurable in advance.

    A corpus is a set of decisions, each measurable before a step is taken.

    How this chapter is built

    M2Substantive

    The derivations are the chapter. A reader who skips the algebra has not learned it.

    basics2/9what the words mean
    concept2/2what to picture
    theory0/2why it works, and when it does not
    mathematics0/9derive it, then compute it
    practice0/6build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    The objective is settled and the architecture is settled. What remains is what the model reads, and that turns out to decide more of the result than either — while being the part most often described rather than measured.

    Accuracy against pretraining set sizeTwo curves against the number of pretraining images on a log scale. The convolutional network is ahead on small datasets. The transformer starts lower, rises faster, and overtakes somewhere around a hundred million images. The crossover is the paper's finding.accuracy against pretraining images1M10M100M1BcrossoverConvNet — prior includedViT — prior learnedAbove the threshold theconstraint costs morethan it returns.The claim was conditional, with a threshold in it. What propagated was "transformers beat CNNs".
    Fig. 3 — Accuracy against pretraining set size. The paper's finding is a crossover with a threshold in it, somewhere near a hundred million images — not a verdict.

    What this chapter covers

    • Corpus construction
    • Data filtering
    • Token budgets
    • Compute
    • Scaling
    • Distributed training

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Bayes' rule 0.PR.04 · Estimators, bias and variance 0.ST.01

    Notation

    • DDataset size in tokens
    • NParameter count
    • BBatch size
    • TSequence length in tokens

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/3 problems0/3 variants0/6 exercisesowes 9 more

    Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.