Md. Asif Uddin

    LectionesPart X

    Reasoning

    Four papers in which nothing about the model changes and the answers get better.

    Sit with how strange the first result is. The weights are frozen. The question is the same. Changing the shape of the prompt so that it shows intermediate steps moves arithmetic accuracy by tens of points.

    The three papers after it are the field working out what that implies. If thinking out loud helps, sample several and vote. If voting helps, search the tree properly. If search needs a signal, supervise the steps rather than the answer.

    By the fourth paper the model is no longer being prompted; it is a subroutine inside an algorithm. That is the transition this part is really about, and it leads directly into Part XII.

    The reading

    1. CoT

      Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

      Wei et al. · NeurIPS · 2022

      Claim
      Prompting with worked examples that show the intermediate steps unlocks arithmetic and symbolic reasoning the same model cannot do when asked for the answer alone.
      Why
      The finding is genuinely odd and worth sitting with, because nothing about the model changed. Note the emergence result: the effect appears only above roughly a hundred billion parameters.
      Read
      Sections 3 and 5.
    2. Self-Consistency

      Self-Consistency Improves Chain of Thought Reasoning in Language Models

      Wang et al. · ICLR 2023 · 2022

      Claim
      Sample several chains of thought and take the majority answer; accuracy rises substantially over greedy decoding.
      Why
      Two pages of idea for large gains, and the beginning of treating inference compute as a dial you can turn.
    3. ToT

      Tree of Thoughts: Deliberate Problem Solving with Large Language Models

      Yao et al. · NeurIPS · 2023

      Claim
      Letting the model propose, evaluate and backtrack over a tree of partial solutions turns generation into search.
      Why
      Where prompting stops being prompting. The Game of 24 result — four per cent to seventy-four — is the number to carry away.
      Read
      Section 3.
    4. Process supervision

      Let's Verify Step by Step

      Lightman et al. · ICLR 2024 · 2023

      Claim
      Supervising each reasoning step beats supervising only the final answer, by a wide margin, on competition mathematics.
      Why
      The clearest statement of process versus outcome supervision, and the reason a modern reasoning model has a verifier attached to it.
      Read
      Sections 3 and 4.