Md. Asif Uddin

    LectionesPart I

    Foundations

    What a deep network is, and the four papers that settled it.

    Every part after this one assumes you can read a network architecture without stopping to ask what a layer is. These four papers are where that vocabulary comes from, and three of them are short.

    Read them in order. They are a single argument delivered over three years: features can be learned rather than designed; scale is the binding constraint; depth is the variable; and depth needs a fix before it works. Each paper is the answer to the problem the previous one exposed.

    There is no attention here and no language. That is deliberate. The Transformer arrives in Part II and it is easier to see what it changed if you have first seen what it replaced.

    The reading

    1. Deep Learning

      Deep Learning

      LeCun, Bengio & Hinton · Nature 521, 436–444 · 2015

      Claim
      Representation learning — features discovered by the machine rather than designed by hand — is what separates deep networks from everything that came before.
      Why
      Written by three of the people who built the field, at the moment it stopped being contested. Twelve pages, no mathematics you cannot follow in your head, and it names every idea the next fifty papers assume you already hold.
      Read
      All of it; it is short. The account of backpropagation is the clearest in the literature.
    2. AlexNet

      ImageNet Classification with Deep Convolutional Neural Networks

      Krizhevsky, Sutskever & Hinton · NeurIPS · 2012

      Claim
      A convolutional network trained on two GPUs cut ImageNet top-5 error from 26% to 15%.
      Why
      The result that started the decade. Nothing in it is new mathematics — ReLU, dropout and augmentation were all known. What was new is the demonstration that the binding constraint had been compute, not ideas.
      Read
      Sections 3 and 4. Skip the GPU-splitting arithmetic; that constraint is gone.
    3. VGG

      Very Deep Convolutional Networks for Large-Scale Image Recognition

      Simonyan & Zisserman · ICLR 2015 · 2014

      Claim
      Stacking small 3x3 convolutions to depth 16 or 19 beats fewer, larger filters at the same computational cost.
      Why
      The paper that isolated depth as the variable. It is also the cleanest ablation table in vision — one thing changes per row — and worth reading as a model of how to argue with experiments rather than adjectives.
      Read
      Tables 1 and 2. The rest is procedure.
    4. ResNet

      Deep Residual Learning for Image Recognition

      He, Zhang, Ren & Sun · CVPR 2016 · 2015

      Claim
      Learning a residual F(x)+x instead of the mapping itself lets a network go to 152 layers and keep improving.
      Why
      The skip connection is in everything you read afterwards, transformers included. Read it for the observation that opens it: deeper plain networks were worse on training error, which rules out overfitting and points squarely at optimisation.
      Read
      Section 3.1 and Figure 1. That figure is the whole argument.