Md. Asif Uddin

    LectionesPart VIII

    Fine-tuning without the compute

    Three papers, five years, and the reason you can tune a large model on one GPU.

    A short part, because the story is short and unusually clean.

    The first paper proposes training a small number of inserted parameters instead of all of them, two years before anyone urgently needed it. The second finds a better shape for those parameters, one that can be folded back into the weights so inference costs nothing extra. The third quantises everything around it so the whole procedure fits on hardware you can rent by the hour.

    Read them in order. LoRA reads as an invention on its own and as a refinement next to adapters, and the second reading is the accurate one.

    The reading

    1. Adapters

      Parameter-Efficient Transfer Learning for NLP

      Houlsby et al. · ICML · 2019

      Claim
      Small bottleneck layers inserted into a frozen model reach full fine-tuning quality while training about 3% of the parameters.
      Why
      Parameter-efficient tuning starts here. Reading it first stops you attributing the idea to LoRA.
      Read
      Section 2 and Figure 2.
    2. LoRA

      LoRA: Low-Rank Adaptation of Large Language Models

      Hu et al. · ICLR 2022 · 2021

      Claim
      The weight update during fine-tuning has low intrinsic rank, so it can be written as the product of two thin matrices — and merged back into the weights at inference for no added latency.
      Why
      The most used method on this list. The zero-latency merge is why it beat adapters, and the rank ablation is the surprise: four is usually enough.
      Read
      Section 4, then Section 7.2 on which rank suffices.
    3. QLoRA

      QLoRA: Efficient Finetuning of Quantized LLMs

      Dettmers, Pagnoni, Holtzman & Zettlemoyer · NeurIPS · 2023

      Claim
      Four-bit NormalFloat quantisation, double quantisation and paged optimisers put 65-billion-parameter fine-tuning on a single 48GB GPU with no measured loss of quality.
      Why
      Three separate engineering ideas, each useful on its own. Section 3 is the one to read: what NF4 is, and why it beats plain four-bit integers on weights that are approximately normal.
      Read
      Section 3.