Chapter V
Instruction Following
From a text predictor to something that answers.
What this chapter covers
- Instruction tuning
- Supervised fine-tuning
- Preference optimisation
- RLHF
- DPO
- Alignment
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.