Chapter 2 V.2
Contrastive Vision-Language Learning
Contrastive pretraining aligns two encoders by making a batch into a classification problem in both directions at once.
A batch becomes a classification problem in both directions at once.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
Two encoders produce vectors in two unrelated spaces. Making those spaces one requires an objective that pulls matched pairs together and pushes everything else apart, and the batch is where the negatives come from.
What this chapter covers
- CLIP
- Image encoder
- Text encoder
- Shared embedding space
- Contrastive objective
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Inner products, norms and cosine similarity 0.LA.03 · Mutual information 0.IT.04 · The softmax Jacobian 0.MC.06
Notation
- τTemperature, in a softmax or a contrastive loss
- BBatch size
- dModel width
- softmaxThe normalised exponential, applied row-wise unless stated
- ∇Gradient operator
- 𝔼Expectation
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.