Md. Asif Uddin

    The commonplace book

    Marginalia

    Notes written in the margins of the work: what I thought of a book, what a model actually does when you use it, and the occasional argument that did not fit anywhere else. Kept in one ledger, numbered as they are written.

    Three parts. The first is the ledger — my opinion of something after reading it. The second is a reading course: what somebody else should read, in what order, and what to take from each one. The third is the shelf those two are worked out on — the textbooks, which are not read in an order and so are not filed in one.

    A living page of marginal notesA reading point moves down a manuscript margin while three annotations, underlines, and cross-references write themselves beside the text.
    Figura IVThe page remembers where it was readThe reader moves down the margin. Notes appear in its wake — one underlined, one circled, one cross-referenced — until the three parts of the commonplace book become one working index.

    RecensionesWhat I thought of it

    1. noteDataset acquisition as a research methodOn asking thirty-seven timesEvery hospital said no. One research group said yes. That is the dataset.
    2. modelEvo 2 (Arc Institute, Nature 2026)What open actually buysThe strongest argument in biology right now for releasing everything.
    3. modelSTATE (Arc Institute, 2025)Two benchmarks, one model, opposite answersBoth claims are in print. Working out how they can both be true is more useful than picking a side.
    4. modelscVI (Lopez et al., Nature Methods 2018)The baseline nobody specifiesThe most useful model in single-cell biology, and the one nobody writes about.
    5. modelViT — An Image is Worth 16x16 Words (Dosovitskiy et al., ICLR 2021)The architecture won, the argument did notThe paper is careful, conditional and honest about its negative result. What spread was a slogan.
    6. modelBoltz-2 (MIT CSAIL and Recursion, 2025)The number a chemist can useA neural network reaching the correlation of a physics simulation, three orders of magnitude faster.
    7. modelAlphaGenome (Google DeepMind, Nature 2026)Refusing the tradeoffThe most convincing large model in genomics right now, and the most frustrating to actually use.
    8. modelGeneformerThe model is fine. The framing wasn'tWhat Geneformer actually is Geneformer is a bidirectional encoder, BERT in shape, pretrained on Genecorpus-30M and later extended to around 95 million cells. The interesting design decision is the input representation. Single-cell expression data is a problem for tokenization. Counts are sparse, noisy, and vary in scale between cells for reasons that have nothing to do with biology. Feeding raw values invites the model to learn sequencing depth. Geneformer's answer is rank value encoding. Within each cell, genes are ranked by expression, but normalized first by that gene's median expression across the entire corpus. A gene that is highly expressed everywhere gets demoted. A gene that is unusually high in this particular cell rises. The result is a ranked list per cell where position encodes something closer to cell identity than to library size. Then the standard masked objective. Hide some genes in the ranking, predict them from the rest.
    9. modelEfficientNet (Tan and Le, 2019 and 2021)EfficientNet, V1 and V2Fewer FLOPs, 2.7x slower, same accuracy. The name did a decade of damage.
    10. modelDINOv3 as a frozen backboneWhat the registers fixedDense features that no longer need apologising for.
    11. modelLlama 3.3 70B and Llama 4 (Meta, 2024 and 2025)Llama 3.3 70B and Llama 4Ten million tokens of advertised context. At a hundred and twenty thousand, 15.6%.
    12. modelSparse MoE, from Jacobs 1991 to Kimi K3Mixture of expertsA parameter count is no longer a capability claim. It is a hardware requirement.
    13. modelKimi K3 and K2.6 (Moonshot AI, 2026)Kimi K3 and K2.6They said the weights would be out by July 27. On July 27, the weights were out.
    14. modelDeepSeek-V4-Pro and V4-Flash (April 2026)DeepSeek V4Everyone quoted the 1.6 trillion. The number that matters is 10%.
    15. bookElements of Causal Inference (Peters, Janzing and Scholkopf, 2017)The Machine Learning Guide to CausalityMIT Press, 2017. It's open access, so the PDF is free, and there are companion Jupyter notebooks with coding exercises. This book connects machine learning with causal reasoning. It shows how algorithms can figure out cause and effect just by looking at regular data.
    16. bookCausal Inference in Statistics: A Primer (Pearl, Glymour and Jewell, 2016)The Starter Guide to Causal InferenceWiley, 2016. Pearl with Madelyn Glymour and Nicholas Jewell, a Berkeley biostatistician, and it shows. This is the only one of the three that behaves like a textbook. This book is a beginner-friendly guide to understanding cause and effect. It uses clear examples and simple math to help you build models and solve real-world problems.
    17. bookCausality: Models, Reasoning, and Inference (Judea Pearl, 2009)Beyond Correlation: A Review of Pearl’s CausalityCambridge University Press, 2nd edition 2009, first published 2000. Eleven chapters, four hundred-odd pages, no attempt to be charming. This book helps AI and statistics move past simple correlation to understand true cause and effect. Pearl provides math tools to help readers map relationships, predict outcomes, and test "what if" scenarios.
    18. bookJudea Pearl & Dana Mackenzie — The Book of WhyThe ladder is the argumentA polemic disguised as a popular science book, and better for it.
    19. modelSAM 3 (Meta, November 2025)SAM 3SAM 1 and 2 answered where. This one answers what, and the clever part is one token.
    20. modelQwen3.8-Max, 2.4T-A95B and 27B (Alibaba, August 2026)Qwen3.8First Max-class Qwen released as weights. Both halves of that headline need checking.
    21. modelThe Qwen series (Alibaba, 2023 onward)Qwen, the familyNot any single model. A ladder you can climb without changing vendor, and weights you can keep.
    22. modelEVA-CLIP-18B (BAAI, 2024)EVA-CLIP, 8B and 18BGoing from 8.1B to 18B bought 0.7 points. The most useful thing in the paper is that number.
    23. modelnnU-Net (Isensee et al., Nature Methods 2021)nnU-Net, the 3D oneIt is not an architecture. That is the whole point, and it is why the paper gets misread.
    24. modelLogistic regression (Berkson 1944, Cox 1958)The instrument, not the competitorEvery linear probe in every self-supervised paper is this model. It is the ruler, and nobody calls it by name.
    25. modelSwin UNETR V2 (NVIDIA, 2023)Swin UNETR V2A transformer that had to be handed convolutions back before it worked properly.
    26. modelConvNeXt V2 (Woo et al., 2023)ConvNeXt V2The least exciting model I use, and it shows up in more of my projects than anything else.
    27. modelDINOv2 (2023) and DINOv3 (2025), Meta AIDINOv2 and DINOv3I still start projects on the three-year-old one. That is not nostalgia, it is licensing.
    28. modelSigLIP 2SigLIP 2Models I actually use, and what annoys me about them.

    LectionesWhat to read, in order

    A reading course in machine learning, published in parts of three to five papers. Every paper appears exactly once, so finishing a part means something. 12 parts so far, 54 papers.

    1. FoundationsWhat a deep network is, and the four papers that settled it.4
    2. The TransformerOne architecture, and the three papers that showed what it could carry.4
    3. AlignmentTurning a model that predicts text into one that answers you.5
    4. Generating before the language modelFour ways to learn a distribution you can sample from.4
    5. Vision after the TransformerWhat happened to computer vision once attention arrived.5
    6. Language and vision in one modelHow a model comes to see and talk about the same thing.5
    7. Position, memory and the cost of attentionAttention is quadratic and the cache does not fit. Five papers about that.5
    8. Fine-tuning without the computeThree papers, five years, and the reason you can tune a large model on one GPU.3
    9. Sparsity and scaleWhy a model can have 671 billion parameters and use 37 billion of them.5
    10. ReasoningFour papers in which nothing about the model changes and the answers get better.4
    11. RetrievalGiving a model access to things it was not trained on.5
    12. AgentsWhere the model stops answering and starts doing.5

    The whole course, with what each part is for

    BibliothecaThe books behind it

    Textbooks rather than papers, and kept apart from the course for a reason: a part of the Lectiones is a week of reading in an order that matters, while a textbook is opened at the chapter you need and closed again. 20 books on 4 shelves.

    The whole shelf, with what each book is for