Md. Asif Uddin

    Chapter 5 IV.5

    Instruction Following and Alignment

    The KL-regularised optimum of a learned reward has a closed form, and inverting it turns preference data directly into a loss.

    A preference dataset defines an implicit reward.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    A pretrained model completes text; it does not answer questions. Turning one into the other means optimising something no dataset directly labels, which is why preference comparisons and a divergence penalty replace a supervised target.

    The presence tokenObject queries used to answer two questions at once: whether the concept is present and where it is. SAM 3 gives presence its own global token, leaves localisation to the queries, and multiplies the two scores.one prompt, two questions"striped cat"presence tokenis the concept here at all?object querieswhere, for each instance×scoreTwo questions that interfered when one set of queries answered both. cgF1 55.7 against a field near 24.
    Fig. 5 — Whether the concept is present and where it is are different questions that used to interfere. Giving presence its own token and multiplying the scores is most of why the numbers moved.

    What this chapter covers

    • Instruction tuning
    • Supervised fine-tuning
    • Preference optimisation
    • RLHF
    • DPO
    • Alignment

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Kullback–Leibler divergence 0.IT.03 · The exponential family 0.PR.06 · Lagrange multipliers 0.OP.04

    Notation

    • βA momentum coefficient in an optimiser; the KL weight in a preference objective
    • σThe logistic function, or a standard deviation
    • softmaxThe normalised exponential, applied row-wise unless stated
    • ∇Gradient operator
    • 𝔼Expectation
    • KLKullback–Leibler divergence

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.