Md. Asif Uddin

    LectionesPart III

    Alignment

    Turning a model that predicts text into one that answers you.

    A base language model does not want to help you. It continues text. The distance between that and an assistant is this part.

    The first paper is the pipeline everyone copied. The four after it are a single sustained argument that the pipeline was more complicated than it needed to be: first the reinforcement learning came out, then the reward model, then the separate stage altogether. Read them in order and you are watching a method get compressed in real time.

    Worth holding onto while you read: nobody has shown that the compression is free. Ask, at each step, what was lost.

    The reading

    1. RLHF

      Training Language Models to Follow Instructions with Human Feedback

      Ouyang et al. · NeurIPS · 2022

      Claim
      A 1.3B model tuned on human preferences is preferred by human raters to a 175B base model.
      Why
      The paper that turned a language model into an assistant. That headline — a hundred times smaller and preferred anyway — is the strongest evidence in the literature that capability and usefulness are separate axes.
      Read
      Section 3: supervised fine-tuning, reward model, PPO. You meet all three again.
    2. DPO

      Direct Preference Optimization: Your Language Model is Secretly a Reward Model

      Rafailov et al. · NeurIPS · 2023

      Claim
      The RLHF objective has a closed form, so preferences can be optimised directly with a classification loss — no reward model, no reinforcement learning.
      Why
      It removed the most fragile component of the pipeline. If you ever implement alignment yourself, this is what you will implement.
      Read
      Section 4. The derivation is short; do it on paper.
    3. Constitutional AI

      Constitutional AI: Harmlessness from AI Feedback

      Bai et al. · Anthropic · 2022

      Claim
      A model can critique and revise its own outputs against a written set of principles, replacing most human harmlessness labels with model feedback.
      Why
      The scaling argument for alignment: human labelling does not keep up with model capability. Also the clearest published account of what a principle is, operationally.
      Read
      Section 3, then the appendix with the constitution itself.
    4. RRHF

      RRHF: Rank Responses to Align Language Models with Human Feedback without tears

      Yuan et al. · NeurIPS · 2023

      Claim
      Ranking sampled responses with a simple loss aligns as well as PPO, at the tuning burden of ordinary fine-tuning.
      Why
      Read it beside DPO. Two groups reaching the same conclusion by different routes — that the reinforcement learning was doing less work than assumed — is stronger evidence than either alone.
    5. ORPO

      ORPO: Monolithic Preference Optimization without Reference Model

      Hong, Lee & Thorne · EMNLP · 2024

      Claim
      Preference alignment folds into supervised fine-tuning as an odds-ratio penalty, with no reference model and no second stage.
      Why
      The end of the compression: three stages became two, then one. A good place to stop and ask what the extra stages had been buying.