LectionesPart IV
Generating before the language model
Four ways to learn a distribution you can sample from.
This part pays back more mathematics per page than any other. Three of the four papers contain a derivation worth doing yourself rather than reading.
The order is chronological and also pedagogical. The variational autoencoder gives you the objective, the GAN gives you an alternative that avoids it, diffusion gives you the objective back in a form that actually trains, and latent diffusion makes it affordable.
If your interest is language, this part is still not optional: the sampling and guidance machinery here reappears in every multimodal system in Part VI.
The reading
- VAE
Auto-Encoding Variational Bayes
Kingma & Welling · ICLR 2014 · 2013
- Claim
- The reparameterisation trick makes the variational bound differentiable, so an approximate posterior can be trained by ordinary gradient descent.
- Why
- The first of the modern generative models. If the evidence lower bound is not yet solid for you, this is where to fix that, and fixing it here saves you three times later.
- Read
- Sections 2.2 and 2.3. Derive the reparameterisation yourself.
- GAN
Generative Adversarial Networks
Goodfellow et al. · NeurIPS · 2014
- Claim
- Training a generator against a discriminator in a minimax game recovers the data distribution at the global optimum.
- Why
- The idea is one page and the proof is two. Read it even though diffusion has largely displaced it — the adversarial framing reappears everywhere afterwards.
- Read
- Section 3 and Section 4.1. Proposition 2 is the paper.
- Diffusion
Denoising Diffusion Probabilistic Models
Ho, Jain & Abbeel · NeurIPS · 2020
- Claim
- A model trained to reverse a fixed noising process, under a simplified loss, matches GAN sample quality.
- Why
- The paper that made diffusion practical. The distance between its variational objective and the three-line training loop is the thing to understand.
- Read
- Algorithms 1 and 2 first, then go back for the derivation.
- Stable Diffusion
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach et al. · CVPR 2022 · 2021
- Claim
- Running diffusion in a compressed latent space rather than pixel space cuts the cost by an order of magnitude without losing fidelity.
- Why
- The engineering paper that put image generation on consumer hardware. The move — do the expensive thing in a smaller space — is worth carrying to problems that have nothing to do with images.