Chapter 7 IV.7
Efficient Adaptation
Adaptation is a rank and a precision decision, and both are bounded by what the update actually has to express.
A rank and a precision decision, both bounded by what the update needs to express.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Full fine-tuning of a seven-billion-parameter model needs more memory than most people have. Two independent ideas — low rank and low precision — make it fit, and each has a cost that can be quantified rather than hoped about.
What this chapter covers
- Fine-tuning
- LoRA
- QLoRA
- Adapters
- Quantisation
- Pruning
- Distillation
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Rank, eigenvalues and the singular value decomposition 0.LA.04 · Floating point 0.NU.01 · Variance and covariance 0.PR.03
Notation
- rRank of a low-rank update
- dModel width
- NParameter count
- WA weight matrix
- VarVariance
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.