LectionesPart X
Reasoning
Four papers in which nothing about the model changes and the answers get better.
Sit with how strange the first result is. The weights are frozen. The question is the same. Changing the shape of the prompt so that it shows intermediate steps moves arithmetic accuracy by tens of points.
The three papers after it are the field working out what that implies. If thinking out loud helps, sample several and vote. If voting helps, search the tree properly. If search needs a signal, supervise the steps rather than the answer.
By the fourth paper the model is no longer being prompted; it is a subroutine inside an algorithm. That is the transition this part is really about, and it leads directly into Part XII.
The reading
- CoT
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei et al. · NeurIPS · 2022
- Claim
- Prompting with worked examples that show the intermediate steps unlocks arithmetic and symbolic reasoning the same model cannot do when asked for the answer alone.
- Why
- The finding is genuinely odd and worth sitting with, because nothing about the model changed. Note the emergence result: the effect appears only above roughly a hundred billion parameters.
- Read
- Sections 3 and 5.
- Self-Consistency
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang et al. · ICLR 2023 · 2022
- Claim
- Sample several chains of thought and take the majority answer; accuracy rises substantially over greedy decoding.
- Why
- Two pages of idea for large gains, and the beginning of treating inference compute as a dial you can turn.
- ToT
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Yao et al. · NeurIPS · 2023
- Claim
- Letting the model propose, evaluate and backtrack over a tree of partial solutions turns generation into search.
- Why
- Where prompting stops being prompting. The Game of 24 result — four per cent to seventy-four — is the number to carry away.
- Read
- Section 3.
- Process supervision
Let's Verify Step by Step
Lightman et al. · ICLR 2024 · 2023
- Claim
- Supervising each reasoning step beats supervising only the final answer, by a wide margin, on competition mathematics.
- Why
- The clearest statement of process versus outcome supervision, and the reason a modern reasoning model has a verifier attached to it.
- Read
- Sections 3 and 4.