Md. Asif Uddin

    Chapter 5 III.5

    Self-Supervised Vision

    Every self-supervised objective must answer why its representations do not collapse to a constant.

    Self-supervision manufactures the label from the input.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Labelled images are scarce and expensive; unlabelled ones are not. An objective built from the image alone is the way out, but the trivial solution — map everything to one vector — has to be excluded by construction.

    Collapse, and what prevents itOn the left, embeddings spread across the space. In the middle, every embedding has fallen onto one point, which drives the agreement objective to zero while carrying no information. On the right, negatives push the embeddings apart and keep the space occupied.agreement onlythe degenerate optimumwith negativestrainloss is zeroand the model is useless"Make two views of thesame photo agree" issolved by a constant.Negatives supply therepulsion. So does astop-gradient, cheaper.
    Fig. 5 — The degenerate optimum: every input mapping to one point drives the agreement objective to zero and carries no information.

    What this chapter covers

    • contrastive learning
    • positive and negative pairs
    • augmentations
    • Siamese learning
    • masked image modelling
    • DINO
    • DINOv2

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Mutual information 0.IT.04 · The softmax Jacobian 0.MC.06 · Inner products, norms and cosine similarity 0.LA.03

    Notation

    • τTemperature, in a softmax or a contrastive loss
    • BBatch size
    • dModel width
    • softmaxThe normalised exponential, applied row-wise unless stated
    • ∇Gradient operator
    • 𝔼Expectation

    Propositions

    1. Prop. 1An objective that only rewards agreement is solved by a constant.Ask a network to map two views of one image to the same vector and the perfect solution is to map everything to the same vector. Every method in this chapter is an answer to that.
    2. Prop. 2The augmentation list is a list of things the model is told not to care about.Two views differ by whatever you applied, and the objective says those differences are noise. Copy a recipe from natural images into a medical domain and you may have declared the finding irrelevant.
    3. Prop. 3An asymmetry between two branches prevents collapse without negatives.Give one branch a predictor and stop the gradient into the other, and the degenerate solution stops being reachable. That is DINO's skeleton, and it is why it needs no negative pairs at all.
    4. Prop. 4Masked image modelling works because pixels are redundant.Hide most of an image and predict the rest. The mask ratio has to be brutal — around three quarters — precisely because neighbouring patches say almost the same thing.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.