Chapter 5 III.5
Self-Supervised Vision
Every self-supervised objective must answer why its representations do not collapse to a constant.
Self-supervision manufactures the label from the input.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Labelled images are scarce and expensive; unlabelled ones are not. An objective built from the image alone is the way out, but the trivial solution — map everything to one vector — has to be excluded by construction.
What this chapter covers
- contrastive learning
- positive and negative pairs
- augmentations
- Siamese learning
- masked image modelling
- DINO
- DINOv2
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Mutual information 0.IT.04 · The softmax Jacobian 0.MC.06 · Inner products, norms and cosine similarity 0.LA.03
Notation
- τTemperature, in a softmax or a contrastive loss
- BBatch size
- dModel width
- softmaxThe normalised exponential, applied row-wise unless stated
- ∇Gradient operator
- 𝔼Expectation
Propositions
- Prop. 1An objective that only rewards agreement is solved by a constant.Ask a network to map two views of one image to the same vector and the perfect solution is to map everything to the same vector. Every method in this chapter is an answer to that.
- Prop. 2The augmentation list is a list of things the model is told not to care about.Two views differ by whatever you applied, and the objective says those differences are noise. Copy a recipe from natural images into a medical domain and you may have declared the finding irrelevant.
- Prop. 3An asymmetry between two branches prevents collapse without negatives.Give one branch a predictor and stop the gradient into the other, and the degenerate solution stops being reachable. That is DINO's skeleton, and it is why it needs no negative pairs at all.
- Prop. 4Masked image modelling works because pixels are redundant.Hide most of an image and predict the rest. The mask ratio has to be brutal — around three quarters — precisely because neighbouring patches say almost the same thing.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.