Md. Asif Uddin

    Book V

    Vision-language (VLM)

    contrastive pretraining, fusion, VLM architectures

    Explain how visual and linguistic representations are joined, and what becomes possible once they share a space.

    7 chapters0 propositions written

    Read first

    vision · large language model llm

    Mathematics assumed — follow these when a step stops making sense.

    Inner products, norms and cosine similarity 0.LA.03 · Mutual information 0.IT.04 · Expectation 0.PR.02

    1. Chapter 1V.1Why Join Vision and LanguageVision and language are joined because the supervision each provides is exactly what the other lacks.Separate modalities · Representation spaces · Multimodality · GroundingM0not yet written
    2. Chapter 2V.2Contrastive Vision-Language LearningContrastive pretraining aligns two encoders by making a batch into a classification problem in both directions at once.CLIP · Image encoder · Text encoder · Shared embedding space · Contrastive objectiveM3not yet written

      0/5 problems0/4 variants0/10 exercisesowes 15 more

    3. Chapter 3V.3Image–Text AlignmentA shared space turns classification into retrieval, but the two modalities do not occupy the same region of it.Paired data · Alignment · Semantic similarity · Zero-shot classificationM2not yet written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    4. Chapter 4V.4Vision-Language GenerationA vision-language model spends context on pixels, and the projector decides how much.Image encoder · Projection layer · Language model · Cross-attention · Visual tokens · Multimodal promptingM2not yet written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    5. Chapter 5V.5VLM ArchitecturesThe architectures differ in where fusion happens, and where it happens decides what can be trained separately.CLIP-style encoders · BLIP and BLIP-2 · LLaVA-style systems · Flamingo-style systems · Modern multimodal LLMsM1not yet written

      0/1 problems0/1 variants0/3 exercisesowes 4 more

    6. Chapter 6V.6Visual ReasoningVisual reasoning fails where the patch grid cannot resolve the question being asked of it.OCR · Grounding · Spatial reasoning · Counting · Chart understanding · Document understanding · Visual question answeringM2not yet written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    7. Chapter 7V.7VLM EvaluationCaption and answer metrics reward surface agreement, so a wrong answer can outscore a right one.VQA · Captioning · Grounding · Hallucination · Multimodal reasoning · RobustnessM2not yet written

      0/3 problems0/3 variants0/6 exercisesowes 9 more

    Problem setWhere the chapters have to be used togetherNot yet written. A book with load-bearing chapters owes at least eight cross-chapter problems and two that reach back into an earlier book.

    Practical connection

    Image–text retrieval, then captioning, then visual question answering.

    Your contrastive loss must agree with the 3x3 similarity matrix worked by hand.

    Verified againstV.2.B02