LectionesPart II
The Transformer
One architecture, and the three papers that showed what it could carry.
Parts VII through XII are all modifications of the architecture in the first paper here. It is worth more of your time than anything else on this list.
Read BERT and GPT-1 as a pair. They were published four months apart, they use the same building block, and they differ in one decision: whether a token may see the tokens after it. Everything that follows — why one model fills in blanks and the other continues text, why one became an encoder and the other became the industry — comes out of that single choice about the mask.
Then GPT-3, which is seventy-five pages long and whose finding is in the first ten.
The reading
- Transformer
Attention Is All You Need
Vaswani et al. · NeurIPS · 2017
- Claim
- Self-attention alone, with no recurrence and no convolution, is sufficient for sequence transduction, and it parallelises.
- Why
- The one paper on this list you should be able to reconstruct from memory. Read it until the shapes are automatic: what Q, K and V are, and why the score is divided by the square root of the key dimension.
- Read
- Section 3 entirely, and work the dimensions in Figure 2 by hand. Section 6.3 can go.
- BERT
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, Chang, Lee & Toutanova · NAACL 2019 · 2018
- Claim
- Pre-training a bidirectional encoder to fill in masked tokens, then fine-tuning it, beats task-specific architectures on eleven tasks at once.
- Why
- The encoder branch of the family tree, and still what you reach for when the task is classification rather than generation.
- Read
- Section 3.1, the masked language model objective. The GLUE tables are history.
- GPT-1
Improving Language Understanding by Generative Pre-Training
Radford, Narasimhan, Salimans & Sutskever · OpenAI technical report · 2018
- Claim
- Generative pre-training on unlabelled text followed by discriminative fine-tuning transfers across tasks with almost no change of architecture per task.
- Why
- Twelve pages, never peer reviewed, and the direct ancestor of everything you use daily. Read it beside BERT and note how modest the claims are.
- Read
- Section 3.
- GPT-3
Language Models are Few-Shot Learners
Brown et al. · NeurIPS · 2020
- Claim
- At 175 billion parameters a language model performs new tasks from examples placed in the prompt, with no gradient update at all.
- Why
- The paper that made in-context learning something everyone had to explain. The broader-impacts section is also worth your time; it was unusual for its date.
- Read
- Sections 1 to 3, then Figures 1.2 and 1.3. Skim the task-by-task appendix.