Chapter IV
Vision-Language Generation
Getting a language model to look.
What this chapter covers
- Image encoder
- Projection layer
- Language model
- Cross-attention
- Visual tokens
- Multimodal prompting
Propositions
Not yet written. The topics above are the plan for this chapter; each will become a proposition with its own figure.