Md. Asif Uddin

    Book III

    Vision

    ViT, DINOv2, self-supervised visual representation

    Explain how a machine represents and interprets visual information, from pixels to self-supervised representations.

    7 chapters26 propositions written

    Read first

    neural networks · foundations

    Mathematics assumed — follow these when a step stops making sense.

    Vectors, matrices and the row-major convention 0.LA.01 · Rank, eigenvalues and the singular value decomposition 0.LA.04 · The chain rule 0.MC.03 · Expectation 0.PR.02 · Confidence intervals 0.ST.02

    1. Chapter 1III.1Images as DataAn image is a tensor whose axes carry meaning, and the spatial structure is information a model can use or throw away.pixels · channels · RGB · resolution · normalisation · convolution · receptive fieldsM23 written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    2. Chapter 2III.2Convolutional ArchitecturesThe CNN lineage is a sequence of answers to one question: how to go deeper without losing the gradient.kernels · convolution · padding · stride · pooling · feature maps · CNN architectures · LeNet · AlexNet · VGG · ResNet · EfficientNetM21 written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    3. Chapter 3III.3Representation LearningTransfer learning is a change of the optimisation's starting point, and that starting point is a prior.low-level features · hierarchical features · transfer learning · pretrained representations · self-supervised learningM14 written

      0/1 problems0/1 variants0/3 exercisesowes 4 more

    4. Chapter 4III.4Vision TransformersOnce an image is a sequence of patches, the attention cost of a picture is the cost of its resolution squared.image patches · patch embeddings · positional information · ViT · hierarchical transformers · Swin TransformerM34 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    5. Chapter 5III.5Self-Supervised VisionEvery self-supervised objective must answer why its representations do not collapse to a constant.contrastive learning · positive and negative pairs · augmentations · Siamese learning · masked image modelling · DINO · DINOv2M34 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    6. Chapter 6III.6Vision TasksClassification, detection and segmentation ask different things of the same backbone, and each asks for its own metric.classification · detection · segmentation · depth · retrievalM25 written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    7. Chapter 7III.7Medical and Scientific VisionAn internal AUROC hides the two things that decide clinical usefulness: where the threshold sits and how rare the disease is.medical imaging · CT · MRI · X-ray · fundus imaging · segmentation · classification · domain shift · clinical evaluation · sensitivity and specificity · precision and recall · F1 · AUROC · calibrationM35 written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    Problem setWhere the chapters have to be used togetherNot yet written. A book with load-bearing chapters owes at least eight cross-chapter problems and two that reach back into an earlier book.

    Practical connection

    A CNN classifier, then a ViT, then a segmentation model.

    Train on a patient split, retrain on a site split, and report the gap as the result.

    Verified againstIII.2.B01 · III.6.B01

    ClosingThe whole book on one plateThe map, the key equations with their glosses, what you should now be able to derive unaided, the vocabulary, the misconceptions and the laboratory. Read it after the chapters, then again before the next book.