Chapter 5 IV.5
Instruction Following and Alignment
The KL-regularised optimum of a learned reward has a closed form, and inverting it turns preference data directly into a loss.
A preference dataset defines an implicit reward.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
A pretrained model completes text; it does not answer questions. Turning one into the other means optimising something no dataset directly labels, which is why preference comparisons and a divergence penalty replace a supervised target.
What this chapter covers
- Instruction tuning
- Supervised fine-tuning
- Preference optimisation
- RLHF
- DPO
- Alignment
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Kullback–Leibler divergence 0.IT.03 · The exponential family 0.PR.06 · Lagrange multipliers 0.OP.04
Notation
- βA momentum coefficient in an optimiser; the KL weight in a preference objective
- σThe logistic function, or a standard deviation
- softmaxThe normalised exponential, applied row-wise unless stated
- ∇Gradient operator
- 𝔼Expectation
- KLKullback–Leibler divergence
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.