LectionesPart V
Vision after the Transformer
What happened to computer vision once attention arrived.
The first paper here makes a claim with a condition attached, and the condition is the interesting part: a plain Transformer beats convolution on images given enough pre-training data, and below roughly a hundred million images it loses.
The next four papers are all, in one way or another, attacks on that condition. Swin puts a useful inductive bias back. DeiT shows the data requirement was an artefact of the training recipe. MAE and DINOv2 remove the labels instead of the data.
Read Part I again if the convolution baselines are hazy. The comparison is the point.
The reading
- ViT
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy et al. · ICLR 2021 · 2020
- Claim
- Cut an image into 16x16 patches, treat them as tokens, and a plain Transformer beats convolutional networks — given enough pre-training data.
- Why
- The condition in that sentence is the paper. Section 4.2 shows where ViT loses, and the honest version of the result is the one with the caveat attached.
- Read
- Sections 3.1 and 4.2. Figure 3 is the one to remember.
- Swin
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu et al. · ICCV · 2021
- Claim
- Computing attention inside shifted local windows gives linear cost in image size and the feature hierarchy a plain ViT lacks.
- Why
- What you actually use for detection and segmentation. Also a clean example of putting a known inductive bias back after it has been removed on principle.
- Read
- Section 3.2, the shifted windowing.
- DeiT
Training data-efficient image transformers and distillation through attention
Touvron et al. · ICML 2021 · 2020
- Claim
- With the right augmentation and a distillation token, a vision Transformer trains on ImageNet alone — no three-hundred-million-image pre-training.
- Why
- It answered ViT's expensive condition within months. Read it as the reply it is.
- MAE
Masked Autoencoders Are Scalable Vision Learners
He et al. · CVPR 2022 · 2021
- Claim
- Mask 75% of the patches, reconstruct them, and you get a strong vision backbone with no labels at all.
- Why
- The asymmetric encoder-decoder is what makes it cheap, and the 75% is not a hyperparameter detail — the paper explains why images tolerate a far higher mask ratio than language does.
- Read
- Section 4 on the mask ratio.
- DINOv2
DINOv2: Learning Robust Visual Features without Supervision
Oquab et al. · TMLR · 2023
- Claim
- Self-supervised features from a curated 142-million-image set transfer without fine-tuning, beating weakly supervised features on most benchmarks.
- Why
- The current answer to which backbone to use. The data-curation pipeline is the part most readers skip and the part that did the work.
- Read
- Section 2, the curation.