LectionesPart I
Foundations
What a deep network is, and the four papers that settled it.
Every part after this one assumes you can read a network architecture without stopping to ask what a layer is. These four papers are where that vocabulary comes from, and three of them are short.
Read them in order. They are a single argument delivered over three years: features can be learned rather than designed; scale is the binding constraint; depth is the variable; and depth needs a fix before it works. Each paper is the answer to the problem the previous one exposed.
There is no attention here and no language. That is deliberate. The Transformer arrives in Part II and it is easier to see what it changed if you have first seen what it replaced.
The reading
- Deep Learning
Deep Learning
LeCun, Bengio & Hinton · Nature 521, 436–444 · 2015
- Claim
- Representation learning — features discovered by the machine rather than designed by hand — is what separates deep networks from everything that came before.
- Why
- Written by three of the people who built the field, at the moment it stopped being contested. Twelve pages, no mathematics you cannot follow in your head, and it names every idea the next fifty papers assume you already hold.
- Read
- All of it; it is short. The account of backpropagation is the clearest in the literature.
- AlexNet
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, Sutskever & Hinton · NeurIPS · 2012
- Claim
- A convolutional network trained on two GPUs cut ImageNet top-5 error from 26% to 15%.
- Why
- The result that started the decade. Nothing in it is new mathematics — ReLU, dropout and augmentation were all known. What was new is the demonstration that the binding constraint had been compute, not ideas.
- Read
- Sections 3 and 4. Skip the GPU-splitting arithmetic; that constraint is gone.
- VGG
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan & Zisserman · ICLR 2015 · 2014
- Claim
- Stacking small 3x3 convolutions to depth 16 or 19 beats fewer, larger filters at the same computational cost.
- Why
- The paper that isolated depth as the variable. It is also the cleanest ablation table in vision — one thing changes per row — and worth reading as a model of how to argue with experiments rather than adjectives.
- Read
- Tables 1 and 2. The rest is procedure.
- ResNet
Deep Residual Learning for Image Recognition
He, Zhang, Ren & Sun · CVPR 2016 · 2015
- Claim
- Learning a residual F(x)+x instead of the mapping itself lets a network go to 152 layers and keep improving.
- Why
- The skip connection is in everything you read afterwards, transformers included. Read it for the observation that opens it: deeper plain networks were worse on training error, which rules out overfitting and points squarely at optimisation.
- Read
- Section 3.1 and Figure 1. That figure is the whole argument.