Chapter 2 III.2
Convolutional Architectures
The CNN lineage is a sequence of answers to one question: how to go deeper without losing the gradient.
The architecture lineage is one question asked five times.
How this chapter is built
M2Substantive
The derivations are the chapter. A reader who skips the algebra has not learned it.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
The mechanics of convolution are settled in Book I. What remains is vision-specific: which arrangements of those mechanics worked, why each generation replaced the last, and what a modern backbone inherits from each.
What this chapter covers
- kernels
- convolution
- padding
- stride
- pooling
- feature maps
- CNN architectures
- LeNet
- AlexNet
- VGG
- ResNet
- EfficientNet
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Inner products, norms and cosine similarity 0.LA.03 · The chain rule 0.MC.03
Notation
- HSpatial height in pixels
- W (spatial)Spatial width in pixels
- CChannels in vision; compute in FLOPs in the scaling chapters
- LNumber of layers
- NParameter count
Propositions
Worked problems
0/3 problems0/3 variants0/6 exercisesowes 9 more
Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.