Md. Asif Uddin

    Chapter 6 III.6

    Vision Tasks

    Classification, detection and segmentation ask different things of the same backbone, and each asks for its own metric.

    The task lives in the head and the loss, not in the backbone.

    How this chapter is built

    M2Substantive

    The derivations are the chapter. A reader who skips the algebra has not learned it.

    basics3/9what the words mean
    concept2/2what to picture
    theory0/2why it works, and when it does not
    mathematics0/9derive it, then compute it
    practice0/6build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    One representation, many questions. What changes between tasks is small and specific — the head, the loss, and the metric — and confusing a metric for a capability is how vision results are most often oversold.

    One image, five tasksA single image and backbone feeding five different heads: classification, detection, segmentation, depth and retrieval. The backbone is the same in each case; what differs is the head and the shape of what comes out.one imageone backbonefive headsencoderunchangedclassificationone labeldetectionboxes + labelssegmentationa label per pixeldeptha distance per pixelretrievala vector to compareNothing about the pixels changed between these five rows. What changed is the question, and with itthe output and the loss that scores it — which is why a good backbone is worth more than a good head.Classification throws the spatial axes away. Everything below it has to keep them, or get them back.
    Fig. 6 — One image and one backbone feeding five heads. What changes is the question, the output shape and the loss.

    What this chapter covers

    • classification
    • detection
    • segmentation
    • depth
    • retrieval

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Confidence intervals 0.ST.02 · Distributions, discrete and continuous 0.PR.01

    Notation

    • HSpatial height in pixels
    • W (spatial)Spatial width in pixels
    • CChannels in vision; compute in FLOPs in the scaling chapters

    Propositions

    1. Prop. 1The task lives in the head and the loss, not in the backbone.One encoder serves classification, detection, segmentation, depth and retrieval. What changes is the shape of the output and the function that scores it.
    2. Prop. 2Detection asks two questions at once and scores them separately.What an object is and where its box sits are different problems with different losses. A prediction can be right about one and wrong about the other, and one number hides which.
    3. Prop. 3A per-pixel task must recover the resolution the encoder discarded.The bottleneck knows what is in the image and no longer knows exactly where. Skip connections are not an optimisation — they are the only place the fine spatial detail still exists.
    4. Prop. 4One photograph cannot fix scale.A small object nearby and a large one far away project to identical pixels. Monocular depth is therefore predicted up to an unknown factor, and a metric claim needs information the image does not contain.
    5. Prop. 5Retrieval learns a geometry, so the categories need not exist yet.A classifier commits to a fixed label set at training time. An embedding commits only to a notion of similarity, and a new category costs one more indexed vector.

    Worked problems

    0/3 problems0/3 variants0/6 exercisesowes 9 more

    Not yet written. At M2 this chapter owes 3 worked problems across 3 distinct variants, and 6 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.