Md. Asif Uddin

Book IV

Vision-language (VLM)

contrastive pretraining, fusion, VLM architectures

Explain how visual and linguistic representations are joined, and what becomes possible once they share a space.

7 chapters0 propositions written

  1. Chapter IWhy Join Vision and Language?Two representation spaces, and the case for one.Separate modalities · Representation spaces · Multimodality · Groundingnot yet written
  2. Chapter IIContrastive Vision-Language LearningCLIP, and the shared embedding space.CLIP · Image encoder · Text encoder · Shared embedding space · Contrastive objectivenot yet written
  3. Chapter IIIImage-Text AlignmentWhat it means for a picture and a sentence to agree.Paired data · Alignment · Semantic similarity · Zero-shot classificationnot yet written
  4. Chapter IVVision-Language GenerationGetting a language model to look.Image encoder · Projection layer · Language model · Cross-attention · Visual tokens · Multimodal promptingnot yet written
  5. Chapter VVLM ArchitecturesThe families, by shape rather than by name.CLIP-style encoders · BLIP and BLIP-2 · LLaVA-style systems · Flamingo-style systems · Modern multimodal LLMsnot yet written
  6. Chapter VIVisual ReasoningWhere looking becomes thinking.OCR · Grounding · Spatial reasoning · Counting · Chart understanding · Document understanding · Visual question answeringnot yet written
  7. Chapter VIIVLM EvaluationMeasuring something that can describe its own mistakes.VQA · Captioning · Grounding · Hallucination · Multimodal reasoning · Robustnessnot yet written