Chapter 5 V.5
VLM Architectures
The architectures differ in where fusion happens, and where it happens decides what can be trained separately.
Where fusion happens decides what can be trained separately.
How this chapter is built
M1Definitional
Mathematics defines, and stops there. At most three displayed equations, no derivations.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
The same two encoders and one decoder can be wired several ways. The choice is not aesthetic: it fixes which components can be frozen, which must be trained jointly, and what a new modality would cost to add.
What this chapter covers
- CLIP-style encoders
- BLIP and BLIP-2
- LLaVA-style systems
- Flamingo-style systems
- Modern multimodal LLMs
Notation
- dModel width
- LNumber of layers
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.
Worked problems
0/1 problems0/1 variants0/3 exercisesowes 4 more
Not yet written. At M1 this chapter owes 1 worked problems across 1 distinct variants, and 3 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.