Md. Asif Uddin

    LectionesPart IV

    Generating before the language model

    Four ways to learn a distribution you can sample from.

    This part pays back more mathematics per page than any other. Three of the four papers contain a derivation worth doing yourself rather than reading.

    The order is chronological and also pedagogical. The variational autoencoder gives you the objective, the GAN gives you an alternative that avoids it, diffusion gives you the objective back in a form that actually trains, and latent diffusion makes it affordable.

    If your interest is language, this part is still not optional: the sampling and guidance machinery here reappears in every multimodal system in Part VI.

    The reading

    1. VAE

      Auto-Encoding Variational Bayes

      Kingma & Welling · ICLR 2014 · 2013

      Claim
      The reparameterisation trick makes the variational bound differentiable, so an approximate posterior can be trained by ordinary gradient descent.
      Why
      The first of the modern generative models. If the evidence lower bound is not yet solid for you, this is where to fix that, and fixing it here saves you three times later.
      Read
      Sections 2.2 and 2.3. Derive the reparameterisation yourself.
    2. GAN

      Generative Adversarial Networks

      Goodfellow et al. · NeurIPS · 2014

      Claim
      Training a generator against a discriminator in a minimax game recovers the data distribution at the global optimum.
      Why
      The idea is one page and the proof is two. Read it even though diffusion has largely displaced it — the adversarial framing reappears everywhere afterwards.
      Read
      Section 3 and Section 4.1. Proposition 2 is the paper.
    3. Diffusion

      Denoising Diffusion Probabilistic Models

      Ho, Jain & Abbeel · NeurIPS · 2020

      Claim
      A model trained to reverse a fixed noising process, under a simplified loss, matches GAN sample quality.
      Why
      The paper that made diffusion practical. The distance between its variational objective and the three-line training loop is the thing to understand.
      Read
      Algorithms 1 and 2 first, then go back for the derivation.
    4. Stable Diffusion

      High-Resolution Image Synthesis with Latent Diffusion Models

      Rombach et al. · CVPR 2022 · 2021

      Claim
      Running diffusion in a compressed latent space rather than pixel space cuts the cost by an order of magnitude without losing fidelity.
      Why
      The engineering paper that put image generation on consumer hardware. The move — do the expensive thing in a smaller space — is worth carrying to problems that have nothing to do with images.