Md. Asif Uddin

    Marginalia II

    Lectiones

    The other half of a commonplace book. Marginalia I is what I thought of something after reading it; this is what somebody else should read, in what order, and what to take from each one.

    Published in parts of three to five papers, because a list of sixty is a thing people bookmark and never open. Every paper appears exactly once. The famous ones are load-bearing in several places at once and it is tempting to repeat them, but a list that repeats itself reads as padding and you lose track of what you have already done. Later parts point back instead.

    12 parts · 54 papers

    1. FoundationsWhat a deep network is, and the four papers that settled it.Deep LearningAlexNetVGGResNet4 papers
    2. The TransformerOne architecture, and the three papers that showed what it could carry.TransformerBERTGPT-1GPT-34 papers
    3. AlignmentTurning a model that predicts text into one that answers you.RLHFDPOConstitutional AIRRHFORPO5 papers
    4. Generating before the language modelFour ways to learn a distribution you can sample from.VAEGANDiffusionStable Diffusion4 papers
    5. Vision after the TransformerWhat happened to computer vision once attention arrived.ViTSwinDeiTMAEDINOv25 papers
    6. Language and vision in one modelHow a model comes to see and talk about the same thing.CLIPBLIPBLIP-2FlamingoLLaVA5 papers
    7. Position, memory and the cost of attentionAttention is quadratic and the cache does not fit. Five papers about that.RoPEFlashAttentionMQAGQARing Attention5 papers
    8. Fine-tuning without the computeThree papers, five years, and the reason you can tune a large model on one GPU.AdaptersLoRAQLoRA3 papers
    9. Sparsity and scaleWhy a model can have 671 billion parameters and use 37 billion of them.Sparse MoESwitchMixtralDeepSeekMoEDeepSeek-V35 papers
    10. ReasoningFour papers in which nothing about the model changes and the answers get better.CoTSelf-ConsistencyToTProcess supervision4 papers
    11. RetrievalGiving a model access to things it was not trained on.RAGREALMRAG surveyGraphRAGCRAG5 papers
    12. AgentsWhere the model stops answering and starts doing.ReActMRKLToolformerGenerative AgentsAutoGen5 papers

    Back to Marginalia I — the reviews