Md. Asif Uddin

    Chapter 4 III.4

    Vision Transformers

    Once an image is a sequence of patches, the attention cost of a picture is the cost of its resolution squared.

    A patch is a token, and that is the entire adaptation.

    How this chapter is built

    M3Load-bearing

    The content is mathematics. Understanding is demonstrated by computation, not recall.

    basics3/11what the words mean
    concept2/2what to picture
    theory0/4why it works, and when it does not
    mathematics0/15derive it, then compute it
    practice0/9build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Attention has no notion of a grid, and an image has nothing else. Cutting the image into patches is the whole bridge — and it imports the quadratic cost of Book II along with the mechanism.

    Image to patches to tokensAn image divided into a grid of fixed-size patches. Each patch is flattened into a vector and multiplied by a learned matrix to give an embedding. A positional embedding is added, and the result is an ordinary sequence of tokens for a transformer encoder.imagepatchesflatten + W+ positionencoderone vectorper patch+p1+p2+p3+p4+p5+p6No vision-specific machinery anywhere in it. Once an image is a sequence of tokens it concatenates with text.
    Fig. 4 — Image to patches to tokens. No vision-specific machinery anywhere in it, which is what made vision-language models straightforward.

    What this chapter covers

    • image patches
    • patch embeddings
    • positional information
    • ViT
    • hierarchical transformers
    • Swin Transformer

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Vectors, matrices and the row-major convention 0.LA.01 · Inner products, norms and cosine similarity 0.LA.03

    Notation

    • PPatch size, in pixels along one side
    • HSpatial height in pixels
    • W (spatial)Spatial width in pixels
    • TSequence length in tokens
    • dModel width
    • CChannels in vision; compute in FLOPs in the scaling chapters
    • hNumber of attention heads

    Propositions

    1. Prop. 1A patch is a token, and that is the entire adaptation.Cut the image into fixed squares, flatten each, project it, add a position. What follows is the transformer encoder from the translation literature, unchanged.
    2. Prop. 2Patch size is the dial between detail and cost, and it is set once.Halving the patch quadruples the tokens and multiplies attention cost sixteenfold. The choice is made at pretraining and everything afterwards inherits it.
    3. Prop. 3Without a positional embedding a ViT sees a bag of patches.Self-attention is permutation-equivariant in two dimensions exactly as it is in one. The grid is not in the operation; it is added to the tokens.
    4. Prop. 4Hierarchy had to be put back before transformers worked on dense tasks.A plain ViT holds one resolution throughout and costs the square of the image. Swin restores the pyramid and confines attention to windows, and the shift between blocks is what stops the windows becoming separate images.

    Worked problems

    0/5 problems0/4 variants0/10 exercisesowes 15 more

    Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.