Chapter 3 IV.3
Pretraining
A pretraining corpus is a set of decisions about duplication, mixture and budget, and each one is measurable in advance.
A corpus is a set of decisions, each measurable before a step is taken.
How this chapter is built
M2Substantive
The derivations are the chapter. A reader who skips the algebra has not learned it.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
The objective is settled and the architecture is settled. What remains is what the model reads, and that turns out to decide more of the result than either — while being the part most often described rather than measured.
What this chapter covers
- Corpus construction
- Data filtering
- Token budgets
- Compute
- Scaling
- Distributed training
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Notation
- DDataset size in tokens
- NParameter count
- BBatch size
- TSequence length in tokens
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/3 problems0/3 variants0/6 exercisesowes 9 more
Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.