Book V
Vision-language (VLM)
contrastive pretraining, fusion, VLM architectures
Explain how visual and linguistic representations are joined, and what becomes possible once they share a space.
7 chapters0 propositions written
Read first
vision · large language model llm
Mathematics assumed — follow these when a step stops making sense.
Inner products, norms and cosine similarity 0.LA.03 · Mutual information 0.IT.04 · Expectation 0.PR.02
- Chapter 1V.1Why Join Vision and LanguageVision and language are joined because the supervision each provides is exactly what the other lacks.Separate modalities · Representation spaces · Multimodality · Grounding
- Chapter 2V.2Contrastive Vision-Language LearningContrastive pretraining aligns two encoders by making a batch into a classification problem in both directions at once.CLIP · Image encoder · Text encoder · Shared embedding space · Contrastive objective
0/5 problems0/4 variants0/10 exercisesowes 15 more
- Chapter 3V.3Image–Text AlignmentA shared space turns classification into retrieval, but the two modalities do not occupy the same region of it.Paired data · Alignment · Semantic similarity · Zero-shot classification
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 4V.4Vision-Language GenerationA vision-language model spends context on pixels, and the projector decides how much.Image encoder · Projection layer · Language model · Cross-attention · Visual tokens · Multimodal prompting
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 5V.5VLM ArchitecturesThe architectures differ in where fusion happens, and where it happens decides what can be trained separately.CLIP-style encoders · BLIP and BLIP-2 · LLaVA-style systems · Flamingo-style systems · Modern multimodal LLMs
0/1 problems0/1 variants0/3 exercisesowes 4 more
- Chapter 6V.6Visual ReasoningVisual reasoning fails where the patch grid cannot resolve the question being asked of it.OCR · Grounding · Spatial reasoning · Counting · Chart understanding · Document understanding · Visual question answering
0/3 problems0/3 variants0/6 exercisesowes 9 more
- Chapter 7V.7VLM EvaluationCaption and answer metrics reward surface agreement, so a wrong answer can outscore a right one.VQA · Captioning · Grounding · Hallucination · Multimodal reasoning · Robustness
0/3 problems0/3 variants0/6 exercisesowes 9 more
Practical connection
Image–text retrieval, then captioning, then visual question answering.
Your contrastive loss must agree with the 3x3 similarity matrix worked by hand.
Verified againstV.2.B02