Md. Asif Uddin

    Chapter 4 V.4

    Vision-Language Generation

    A vision-language model spends context on pixels, and the projector decides how much.

    A vision-language model spends context on pixels.

    How this chapter is built

    M2Substantive

    The derivations are the chapter. A reader who skips the algebra has not learned it.

    basics2/9what the words mean
    concept2/2what to picture
    theory0/2why it works, and when it does not
    mathematics0/9derive it, then compute it
    practice0/6build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    Retrieval only ranks; generation must produce language conditioned on an image. Feeding an image into a language model means turning it into tokens, and every visual token is a text token that no longer fits.

    Fewer operations, more time, same accuracyTwo models that both reach 84.0% on ImageNet. One uses 1.8 times fewer floating-point operations and runs 2.7 times slower on the same hardware, because depthwise convolutions are bound by memory bandwidth rather than arithmetic.both land at 84.0% on ImageNetEfficientNet-B6FLOPsarithmetictimewall clockResNet-RS-350FLOPsarithmetictimewall clockAccelerators are bound by memory bandwidth, not arithmetic. "Efficient" meant FLOP-efficient, and a decadeof practitioners read it as fast.
    Fig. 4 — Two models at the same accuracy: one with 1.8x fewer operations and 2.7x more wall-clock time. Accelerators are bound by memory bandwidth, not arithmetic.

    What this chapter covers

    • Image encoder
    • Projection layer
    • Language model
    • Cross-attention
    • Visual tokens
    • Multimodal prompting

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Vectors, matrices and the row-major convention 0.LA.01

    Notation

    • TSequence length in tokens
    • dModel width
    • PPatch size, in pixels along one side
    • LNumber of layers

    Propositions

    Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.

    Worked problems

    0/3 problems0/3 variants0/6 exercisesowes 9 more

    Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.