Marginalia II
Lectiones
The other half of a commonplace book. Marginalia I is what I thought of something after reading it; this is what somebody else should read, in what order, and what to take from each one.
Published in parts of three to five papers, because a list of sixty is a thing people bookmark and never open. Every paper appears exactly once. The famous ones are load-bearing in several places at once and it is tempting to repeat them, but a list that repeats itself reads as padding and you lose track of what you have already done. Later parts point back instead.
12 parts · 54 papers
- FoundationsWhat a deep network is, and the four papers that settled it.Deep LearningAlexNetVGGResNet4 papers
- The TransformerOne architecture, and the three papers that showed what it could carry.TransformerBERTGPT-1GPT-34 papers
- AlignmentTurning a model that predicts text into one that answers you.RLHFDPOConstitutional AIRRHFORPO5 papers
- Generating before the language modelFour ways to learn a distribution you can sample from.VAEGANDiffusionStable Diffusion4 papers
- Vision after the TransformerWhat happened to computer vision once attention arrived.ViTSwinDeiTMAEDINOv25 papers
- Language and vision in one modelHow a model comes to see and talk about the same thing.CLIPBLIPBLIP-2FlamingoLLaVA5 papers
- Position, memory and the cost of attentionAttention is quadratic and the cache does not fit. Five papers about that.RoPEFlashAttentionMQAGQARing Attention5 papers
- Fine-tuning without the computeThree papers, five years, and the reason you can tune a large model on one GPU.AdaptersLoRAQLoRA3 papers
- Sparsity and scaleWhy a model can have 671 billion parameters and use 37 billion of them.Sparse MoESwitchMixtralDeepSeekMoEDeepSeek-V35 papers
- ReasoningFour papers in which nothing about the model changes and the answers get better.CoTSelf-ConsistencyToTProcess supervision4 papers
- RetrievalGiving a model access to things it was not trained on.RAGREALMRAG surveyGraphRAGCRAG5 papers
- AgentsWhere the model stops answering and starts doing.ReActMRKLToolformerGenerative AgentsAutoGen5 papers