Book III
Vision
ViT, DINOv2, self-supervised visual representation
Explain how a machine represents and interprets visual information, from pixels to self-supervised representations.
7 chapters26 propositions written
Read first
Mathematics assumed — follow these when a step stops making sense.
Vectors, matrices and the row-major convention 0.LA.01 · Rank, eigenvalues and the singular value decomposition 0.LA.04 · The chain rule 0.MC.03 · Expectation 0.PR.02 · Confidence intervals 0.ST.02
- Chapter 1III.1Images as DataAn image is a tensor whose axes carry meaning, and the spatial structure is information a model can use or throw away.pixels · channels · RGB · resolution · normalisation · convolution · receptive fields
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 2III.2Convolutional ArchitecturesThe CNN lineage is a sequence of answers to one question: how to go deeper without losing the gradient.kernels · convolution · padding · stride · pooling · feature maps · CNN architectures · LeNet · AlexNet · VGG · ResNet · EfficientNet
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 3III.3Representation LearningTransfer learning is a change of the optimisation's starting point, and that starting point is a prior.low-level features · hierarchical features · transfer learning · pretrained representations · self-supervised learning
0/1 problems0/1 variants0/3 exercisesowes 4 more
- Chapter 4III.4Vision TransformersOnce an image is a sequence of patches, the attention cost of a picture is the cost of its resolution squared.image patches · patch embeddings · positional information · ViT · hierarchical transformers · Swin Transformer
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 5III.5Self-Supervised VisionEvery self-supervised objective must answer why its representations do not collapse to a constant.contrastive learning · positive and negative pairs · augmentations · Siamese learning · masked image modelling · DINO · DINOv2
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 6III.6Vision TasksClassification, detection and segmentation ask different things of the same backbone, and each asks for its own metric.classification · detection · segmentation · depth · retrieval
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 7III.7Medical and Scientific VisionAn internal AUROC hides the two things that decide clinical usefulness: where the threshold sits and how rare the disease is.medical imaging · CT · MRI · X-ray · fundus imaging · segmentation · classification · domain shift · clinical evaluation · sensitivity and specificity · precision and recall · F1 · AUROC · calibration
0/5 problems0/4 variants0/10 exercisesowes 15 more
Practical connection
A CNN classifier, then a ViT, then a segmentation model.
Train on a patient split, retrain on a site split, and report the gap as the result.
Verified againstIII.2.B01 · III.6.B01