Md. Asif Uddin

Chapter IV

Vision-Language Generation

Getting a language model to look.

What this chapter covers

  • Image encoder
  • Projection layer
  • Language model
  • Cross-attention
  • Visual tokens
  • Multimodal prompting

Propositions

Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.