Md. Asif Uddin

    Chapter 7 IV.7

    Efficient Adaptation

    Adaptation is a rank and a precision decision, and both are bounded by what the update actually has to express.

    A rank and a precision decision, both bounded by what the update needs to express.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Full fine-tuning of a seven-billion-parameter model needs more memory than most people have. Two independent ideas — low rank and low precision — make it fit, and each has a cost that can be quantified rather than hoped about.

    Total parameters against active parametersA router scores one token against many expert feed-forward layers and runs only the top few. Every expert must be stored in memory; only the fired ones do arithmetic. Compute tracks the small number, memory the large one.one tokentokenroutertop-kevery expert stored — memorytwo fired — computeA parameter count isno longer a capabilityclaim. It is a hardwarerequirement.
    Fig. 7 — Every expert is stored; only the routed few compute. Memory tracks the big number and arithmetic tracks the small one, which is the whole reason anyone does this.

    What this chapter covers

    • Fine-tuning
    • LoRA
    • QLoRA
    • Adapters
    • Quantisation
    • Pruning
    • Distillation

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Rank, eigenvalues and the singular value decomposition 0.LA.04 · Floating point 0.NU.01 · Variance and covariance 0.PR.03

    Notation

    • rRank of a low-rank update
    • dModel width
    • NParameter count
    • WA weight matrix
    • VarVariance

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.