Md. Asif Uddin

    LectionesPart V

    Vision after the Transformer

    What happened to computer vision once attention arrived.

    The first paper here makes a claim with a condition attached, and the condition is the interesting part: a plain Transformer beats convolution on images given enough pre-training data, and below roughly a hundred million images it loses.

    The next four papers are all, in one way or another, attacks on that condition. Swin puts a useful inductive bias back. DeiT shows the data requirement was an artefact of the training recipe. MAE and DINOv2 remove the labels instead of the data.

    Read Part I again if the convolution baselines are hazy. The comparison is the point.

    The reading

    1. ViT

      An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

      Dosovitskiy et al. · ICLR 2021 · 2020

      Claim
      Cut an image into 16x16 patches, treat them as tokens, and a plain Transformer beats convolutional networks — given enough pre-training data.
      Why
      The condition in that sentence is the paper. Section 4.2 shows where ViT loses, and the honest version of the result is the one with the caveat attached.
      Read
      Sections 3.1 and 4.2. Figure 3 is the one to remember.
    2. Swin

      Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

      Liu et al. · ICCV · 2021

      Claim
      Computing attention inside shifted local windows gives linear cost in image size and the feature hierarchy a plain ViT lacks.
      Why
      What you actually use for detection and segmentation. Also a clean example of putting a known inductive bias back after it has been removed on principle.
      Read
      Section 3.2, the shifted windowing.
    3. DeiT

      Training data-efficient image transformers and distillation through attention

      Touvron et al. · ICML 2021 · 2020

      Claim
      With the right augmentation and a distillation token, a vision Transformer trains on ImageNet alone — no three-hundred-million-image pre-training.
      Why
      It answered ViT's expensive condition within months. Read it as the reply it is.
    4. MAE

      Masked Autoencoders Are Scalable Vision Learners

      He et al. · CVPR 2022 · 2021

      Claim
      Mask 75% of the patches, reconstruct them, and you get a strong vision backbone with no labels at all.
      Why
      The asymmetric encoder-decoder is what makes it cheap, and the 75% is not a hyperparameter detail — the paper explains why images tolerate a far higher mask ratio than language does.
      Read
      Section 4 on the mask ratio.
    5. DINOv2

      DINOv2: Learning Robust Visual Features without Supervision

      Oquab et al. · TMLR · 2023

      Claim
      Self-supervised features from a curated 142-million-image set transfer without fine-tuning, beating weakly supervised features on most benchmarks.
      Why
      The current answer to which backbone to use. The data-curation pipeline is the part most readers skip and the part that did the work.
      Read
      Section 2, the curation.