Chapter 6 III.6
Vision Tasks
Classification, detection and segmentation ask different things of the same backbone, and each asks for its own metric.
The task lives in the head and the loss, not in the backbone.
How this chapter is built
M2Substantive
The derivations are the chapter. A reader who skips the algebra has not learned it.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
One representation, many questions. What changes between tasks is small and specific — the head, the loss, and the metric — and confusing a metric for a capability is how vision results are most often oversold.
What this chapter covers
- classification
- detection
- segmentation
- depth
- retrieval
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Confidence intervals 0.ST.02 · Distributions, discrete and continuous 0.PR.01
Notation
- HSpatial height in pixels
- W (spatial)Spatial width in pixels
- CChannels in vision; compute in FLOPs in the scaling chapters
Propositions
- Prop. 1The task lives in the head and the loss, not in the backbone.One encoder serves classification, detection, segmentation, depth and retrieval. What changes is the shape of the output and the function that scores it.
- Prop. 2Detection asks two questions at once and scores them separately.What an object is and where its box sits are different problems with different losses. A prediction can be right about one and wrong about the other, and one number hides which.
- Prop. 3A per-pixel task must recover the resolution the encoder discarded.The bottleneck knows what is in the image and no longer knows exactly where. Skip connections are not an optimisation — they are the only place the fine spatial detail still exists.
- Prop. 4One photograph cannot fix scale.A small object nearby and a large one far away project to identical pixels. Monocular depth is therefore predicted up to an unknown factor, and a metric claim needs information the image does not contain.
- Prop. 5Retrieval learns a geometry, so the categories need not exist yet.A classifier commits to a fixed label set at training time. An embedding commits only to a notion of similarity, and a new category costs one more indexed vector.
Worked problems
0/3 problems0/3 variants0/6 exercisesowes 9 more
Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.