Md. Asif Uddin

    LectionesPart VI

    Language and vision in one model

    How a model comes to see and talk about the same thing.

    One idea runs through all five: put images and text into a space where distance means the same thing for both.

    CLIP builds that space. The four papers after it are increasingly cheap ways to reuse it — from training the bridge, to freezing both towers and training only the connector, to LLaVA, which is a linear projection and some generated data and which you can reproduce on hardware you own.

    The trend is worth naming while you read. Over three years the trainable fraction of these systems falls by two orders of magnitude and the performance goes up.

    The reading

    1. CLIP

      Learning Transferable Visual Models From Natural Language Supervision

      Radford et al. · ICML · 2021

      Claim
      Contrastive training on 400 million image-text pairs gives a model that classifies images it was never trained to classify, from a written description alone.
      Why
      The joint embedding space underneath most multimodal work, and still the default retrieval model. Section 3.1.4 on prompt engineering is the first appearance of a problem that later swallowed the field.
      Read
      Sections 2.1 to 2.4, and the pseudocode in Figure 3.
    2. BLIP

      BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

      Li, Li, Xiong & Hoi · ICML · 2022

      Claim
      One model for both understanding and generation, trained on web data that a captioner-and-filter pair, bootstrapped from the model itself, has cleaned.
      Why
      The captioning-and-filtering loop is the transferable idea: the model improves its own training data. That move is now standard and this is where to see it stated plainly.
      Read
      Section 3.3.
    3. BLIP-2

      BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

      Li, Li, Savarese & Hoi · ICML · 2023

      Claim
      A small trained bridge — the Q-Former — connects a frozen image encoder to a frozen language model, matching Flamingo with 54 times fewer trainable parameters.
      Why
      The pattern almost every multimodal system now follows: freeze both towers, train the adapter between them.
      Read
      Section 3.1, the two-stage training of the Q-Former.
    4. Flamingo

      Flamingo: a Visual Language Model for Few-Shot Learning

      Alayrac et al. · NeurIPS · 2022

      Claim
      Gated cross-attention layers inserted into a frozen language model give few-shot learning over interleaved sequences of images and text.
      Why
      The first convincing few-shot vision-language model. The gating is the part to understand: it is why you can insert new layers into a trained network without destroying what it knows.
      Read
      Section 2.2.
    5. LLaVA

      Visual Instruction Tuning

      Liu, Li, Wu & Lee · NeurIPS · 2023

      Claim
      Instruction data generated by a language model from captions and bounding boxes is enough to turn a language model into a visual assistant, through a single linear projection.
      Why
      The cheapest working recipe on this list and the one you can actually reproduce. The contribution is the data, not the model, which is worth noticing.
      Read
      Section 3, how the instruction data was made.