Md. Asif Uddin

    Chapter 3 III.3

    Representation Learning

    Transfer learning is a change of the optimisation's starting point, and that starting point is a prior.

    The hierarchy is learned, not designed.

    How this chapter is built

    M1Definitional

    Mathematics defines, and stops there. At most three displayed equations, no derivations.

    basics2/6what the words mean
    concept2/2what to picture
    theory0/1why it works, and when it does not
    mathematics0/4derive it, then compute it
    practice0/5build it, break it, read the papers

    Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.

    Before you start

    The problem

    A trained backbone is worth more than its accuracy: its intermediate features transfer. Understanding what transfers, and how much of the network may move, is what makes a small dataset workable at all.

    What each depth responds toFour stages of a convolutional stack. Early layers respond to edges and colour, middle layers to texture and parts, late layers to whole objects. The early ones are generic across tasks, which is what makes transfer work.genericspecificlayer 1edges and colour blobslayer 2corners, texture, repeatslayer 3parts — an eye, a wheellayer 4objects and whole scenesNobody specified this. The hierarchy is what gradient descent produces with the loss at the far end.An edge is an edge whatever the task, which is why early layers transfer and late ones do not.
    Fig. 3 — What each depth responds to, from edges to whole objects. Nobody specified it; it is what the loss at the far end produces.

    What this chapter covers

    • low-level features
    • hierarchical features
    • transfer learning
    • pretrained representations
    • self-supervised learning

    Apparatus

    The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.

    Gradient descent 0.OP.02

    Notation

    • θAll parameters of a model, taken together
    • dModel width

    Propositions

    1. Prop. 1The hierarchy is learned, not designed.Early layers converge on edges and colour, middle layers on texture and parts, late layers on objects. Nobody specified that ordering; it is what gradient descent produces when the loss sits at the far end.
    2. Prop. 2How much of the network you let move is a statement about how much data you have.Freezing everything and training a linear head needs hundreds of labels. Fine-tuning everything needs tens of thousands. The choice is not a preference; it is arithmetic about capacity and examples.
    3. Prop. 3Pretraining changes where the search starts, not what it can reach.The architecture fixes the family of functions. Pretrained weights are a point inside it that is already useful, which matters most exactly where labels are scarce.
    4. Prop. 4Self-supervision manufactures the label from the input.Corrupt the image in a way you control, and the answer is known without an annotator. What the model must learn in order to undo the corruption is the representation you were after.

    Worked problems

    0/1 problems0/1 variants0/3 exercisesowes 4 more

    Not yet written. At M1 this chapter owes 1 worked problems across 1 distinct variants, and 3 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.