Md. Asif Uddin

    Chapter 2 III.2

    Convolutional Architectures

    The CNN lineage is a sequence of answers to one question: how to go deeper without losing the gradient.

    The architecture lineage is one question asked five times.

    How this chapter is built

    M2Substantive

    The derivations are the chapter. A reader who skips the algebra has not learned it.

    basics3/9what the words mean
    concept2/2what to picture
    theory0/2why it works, and when it does not
    mathematics0/9derive it, then compute it
    practice0/6build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    The mechanics of convolution are settled in Book I. What remains is vision-specific: which arrangements of those mechanics worked, why each generation replaced the last, and what a modern backbone inherits from each.

    Five architectures, one questionFive convolutional architectures in order, with their depth in layers. Each is an answer to the same problem: how to add depth without the optimisation failing. ResNet's residual connection is the step that made depth cheap.how do you get deeper?5 layersLeNet1998it works at all8 layersAlexNet2012ReLU, dropout, GPUs19 layersVGG2014only 3×3, stacked152 layersResNet2015an additive path home66 layersEfficientNet2019scale the three axes togetherDepth axis is logarithmic. Before ResNet the fight was optimisation; after it, tuning.
    Fig. 2 — Five architectures asking one question: how to get deeper without the optimisation failing. The depth axis is logarithmic.

    What this chapter covers

    • kernels
    • convolution
    • padding
    • stride
    • pooling
    • feature maps
    • CNN architectures
    • LeNet
    • AlexNet
    • VGG
    • ResNet
    • EfficientNet

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Inner products, norms and cosine similarity 0.LA.03 · The chain rule 0.MC.03

    Notation

    • HSpatial height in pixels
    • W (spatial)Spatial width in pixels
    • CChannels in vision; compute in FLOPs in the scaling chapters
    • LNumber of layers
    • NParameter count

    Propositions

    1. Prop. 1The architecture lineage is one question asked five times.LeNet through EfficientNet is not a parade of unrelated ideas. Each is an answer to how you add depth without the optimisation falling over, and the residual connection is the one that settled it.

    Worked problems

    0/3 problems0/3 variants0/6 exercisesowes 9 more

    Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.