LectionesPart VI
Language and vision in one model
How a model comes to see and talk about the same thing.
One idea runs through all five: put images and text into a space where distance means the same thing for both.
CLIP builds that space. The four papers after it are increasingly cheap ways to reuse it — from training the bridge, to freezing both towers and training only the connector, to LLaVA, which is a linear projection and some generated data and which you can reproduce on hardware you own.
The trend is worth naming while you read. Over three years the trainable fraction of these systems falls by two orders of magnitude and the performance goes up.
The reading
- CLIP
Learning Transferable Visual Models From Natural Language Supervision
Radford et al. · ICML · 2021
- Claim
- Contrastive training on 400 million image-text pairs gives a model that classifies images it was never trained to classify, from a written description alone.
- Why
- The joint embedding space underneath most multimodal work, and still the default retrieval model. Section 3.1.4 on prompt engineering is the first appearance of a problem that later swallowed the field.
- Read
- Sections 2.1 to 2.4, and the pseudocode in Figure 3.
- BLIP
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Li, Li, Xiong & Hoi · ICML · 2022
- Claim
- One model for both understanding and generation, trained on web data that a captioner-and-filter pair, bootstrapped from the model itself, has cleaned.
- Why
- The captioning-and-filtering loop is the transferable idea: the model improves its own training data. That move is now standard and this is where to see it stated plainly.
- Read
- Section 3.3.
- BLIP-2
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Li, Li, Savarese & Hoi · ICML · 2023
- Claim
- A small trained bridge — the Q-Former — connects a frozen image encoder to a frozen language model, matching Flamingo with 54 times fewer trainable parameters.
- Why
- The pattern almost every multimodal system now follows: freeze both towers, train the adapter between them.
- Read
- Section 3.1, the two-stage training of the Q-Former.
- Flamingo
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac et al. · NeurIPS · 2022
- Claim
- Gated cross-attention layers inserted into a frozen language model give few-shot learning over interleaved sequences of images and text.
- Why
- The first convincing few-shot vision-language model. The gating is the part to understand: it is why you can insert new layers into a trained network without destroying what it knows.
- Read
- Section 2.2.
- LLaVA
Visual Instruction Tuning
Liu, Li, Wu & Lee · NeurIPS · 2023
- Claim
- Instruction data generated by a language model from captions and bounding boxes is enough to turn a language model into a visual assistant, through a single linear projection.
- Why
- The cheapest working recipe on this list and the one you can actually reproduce. The contribution is the data, not the model, which is worth noticing.
- Read
- Section 3, how the instruction data was made.