Md. Asif Uddin

A course in deep learning

Elementa

What a transformer is doing, what a vision model sees, and how the two are joined into a vision-language system. Read in order, like Euclid. Each entry is a proposition rather than a post: it states a claim, names the propositions it depends on, and carries one figure. If the figure cannot be drawn, the proposition is not ready to be written.

Book IFoundations

tokenisation, attention, the transformer block

Book IIVision

ViT, DINOv2, self-supervised visual representation

Not yet written.

Book IIILarge Language Model (LLM)

Transformers, attention mechanisms, prompt engineering, fine-tuning.

Not yet written.

Book IVVision-language (VLM)

contrastive pretraining, fusion, VLM architectures

Not yet written.

Book VCausal inference

DAGs, interventions, counterfactuals, perturbation

Not yet written.

Book VIBioinformatics

single-cell data, gene regulatory networks, perturbation screens

Not yet written.

Book VIIPractice

losses, training dynamics, evaluation

Not yet written.