Book IV
Vision-language (VLM)
contrastive pretraining, fusion, VLM architectures
Explain how visual and linguistic representations are joined, and what becomes possible once they share a space.
7 chapters0 propositions written
- Chapter IWhy Join Vision and Language?Two representation spaces, and the case for one.Separate modalities · Representation spaces · Multimodality · Groundingnot yet written
- Chapter IIContrastive Vision-Language LearningCLIP, and the shared embedding space.CLIP · Image encoder · Text encoder · Shared embedding space · Contrastive objectivenot yet written
- Chapter IIIImage-Text AlignmentWhat it means for a picture and a sentence to agree.Paired data · Alignment · Semantic similarity · Zero-shot classificationnot yet written
- Chapter IVVision-Language GenerationGetting a language model to look.Image encoder · Projection layer · Language model · Cross-attention · Visual tokens · Multimodal promptingnot yet written
- Chapter VVLM ArchitecturesThe families, by shape rather than by name.CLIP-style encoders · BLIP and BLIP-2 · LLaVA-style systems · Flamingo-style systems · Modern multimodal LLMsnot yet written
- Chapter VIVisual ReasoningWhere looking becomes thinking.OCR · Grounding · Spatial reasoning · Counting · Chart understanding · Document understanding · Visual question answeringnot yet written
- Chapter VIIVLM EvaluationMeasuring something that can describe its own mistakes.VQA · Captioning · Grounding · Hallucination · Multimodal reasoning · Robustnessnot yet written