A course in deep learning
Elementa
What a transformer is doing, what a vision model sees, and how the two are joined into a vision-language system. Read in order, like Euclid. Each entry is a proposition rather than a post: it states a claim, names the propositions it depends on, and carries one figure. If the figure cannot be drawn, the proposition is not ready to be written.
Book IFoundations
tokenisation, attention, the transformer block
Book IIVision
ViT, DINOv2, self-supervised visual representation
Not yet written.
Book IIILarge Language Model (LLM)
Transformers, attention mechanisms, prompt engineering, fine-tuning.
Not yet written.
Book IVVision-language (VLM)
contrastive pretraining, fusion, VLM architectures
Not yet written.
Book VCausal inference
DAGs, interventions, counterfactuals, perturbation
Not yet written.
Book VIBioinformatics
single-cell data, gene regulatory networks, perturbation screens
Not yet written.
Book VIIPractice
losses, training dynamics, evaluation
Not yet written.