Md. Asif Uddin

    LectionesPart II

    The Transformer

    One architecture, and the three papers that showed what it could carry.

    Parts VII through XII are all modifications of the architecture in the first paper here. It is worth more of your time than anything else on this list.

    Read BERT and GPT-1 as a pair. They were published four months apart, they use the same building block, and they differ in one decision: whether a token may see the tokens after it. Everything that follows — why one model fills in blanks and the other continues text, why one became an encoder and the other became the industry — comes out of that single choice about the mask.

    Then GPT-3, which is seventy-five pages long and whose finding is in the first ten.

    The reading

    1. Transformer

      Attention Is All You Need

      Vaswani et al. · NeurIPS · 2017

      Claim
      Self-attention alone, with no recurrence and no convolution, is sufficient for sequence transduction, and it parallelises.
      Why
      The one paper on this list you should be able to reconstruct from memory. Read it until the shapes are automatic: what Q, K and V are, and why the score is divided by the square root of the key dimension.
      Read
      Section 3 entirely, and work the dimensions in Figure 2 by hand. Section 6.3 can go.
    2. BERT

      BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

      Devlin, Chang, Lee & Toutanova · NAACL 2019 · 2018

      Claim
      Pre-training a bidirectional encoder to fill in masked tokens, then fine-tuning it, beats task-specific architectures on eleven tasks at once.
      Why
      The encoder branch of the family tree, and still what you reach for when the task is classification rather than generation.
      Read
      Section 3.1, the masked language model objective. The GLUE tables are history.
    3. GPT-1

      Improving Language Understanding by Generative Pre-Training

      Radford, Narasimhan, Salimans & Sutskever · OpenAI technical report · 2018

      Claim
      Generative pre-training on unlabelled text followed by discriminative fine-tuning transfers across tasks with almost no change of architecture per task.
      Why
      Twelve pages, never peer reviewed, and the direct ancestor of everything you use daily. Read it beside BERT and note how modest the claims are.
      Read
      Section 3.
    4. GPT-3

      Language Models are Few-Shot Learners

      Brown et al. · NeurIPS · 2020

      Claim
      At 175 billion parameters a language model performs new tasks from examples placed in the prompt, with no gradient update at all.
      Why
      The paper that made in-context learning something everyone had to explain. The broader-impacts section is also worth your time; it was unusual for its date.
      Read
      Sections 1 to 3, then Figures 1.2 and 1.3. Skim the task-by-task appendix.