LectionesPart III
Alignment
Turning a model that predicts text into one that answers you.
A base language model does not want to help you. It continues text. The distance between that and an assistant is this part.
The first paper is the pipeline everyone copied. The four after it are a single sustained argument that the pipeline was more complicated than it needed to be: first the reinforcement learning came out, then the reward model, then the separate stage altogether. Read them in order and you are watching a method get compressed in real time.
Worth holding onto while you read: nobody has shown that the compression is free. Ask, at each step, what was lost.
The reading
- RLHF
Training Language Models to Follow Instructions with Human Feedback
Ouyang et al. · NeurIPS · 2022
- Claim
- A 1.3B model tuned on human preferences is preferred by human raters to a 175B base model.
- Why
- The paper that turned a language model into an assistant. That headline — a hundred times smaller and preferred anyway — is the strongest evidence in the literature that capability and usefulness are separate axes.
- Read
- Section 3: supervised fine-tuning, reward model, PPO. You meet all three again.
- DPO
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov et al. · NeurIPS · 2023
- Claim
- The RLHF objective has a closed form, so preferences can be optimised directly with a classification loss — no reward model, no reinforcement learning.
- Why
- It removed the most fragile component of the pipeline. If you ever implement alignment yourself, this is what you will implement.
- Read
- Section 4. The derivation is short; do it on paper.
- Constitutional AI
Constitutional AI: Harmlessness from AI Feedback
Bai et al. · Anthropic · 2022
- Claim
- A model can critique and revise its own outputs against a written set of principles, replacing most human harmlessness labels with model feedback.
- Why
- The scaling argument for alignment: human labelling does not keep up with model capability. Also the clearest published account of what a principle is, operationally.
- Read
- Section 3, then the appendix with the constitution itself.
- RRHF
RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Yuan et al. · NeurIPS · 2023
- Claim
- Ranking sampled responses with a simple loss aligns as well as PPO, at the tuning burden of ordinary fine-tuning.
- Why
- Read it beside DPO. Two groups reaching the same conclusion by different routes — that the reinforcement learning was doing less work than assumed — is stronger evidence than either alone.
- ORPO
ORPO: Monolithic Preference Optimization without Reference Model
Hong, Lee & Thorne · EMNLP · 2024
- Claim
- Preference alignment folds into supervised fine-tuning as an odds-ratio penalty, with no reference model and no second stage.
- Why
- The end of the compression: three stages became two, then one. A good place to stop and ask what the extra stages had been buying.