Md. Asif Uddin

    Elementa

    The figure library

    Every diagram on this site, at its own address. All are hand-written inline SVG using the site palette, so they stay legible in both themes and paste into a slide deck without becoming a screenshot of a screenshot.

    FigPipelineHierarchiRetina

    Plate I — HierarchiRetinapermalink
    The HierarchiRetina three-stage pipelineA fundus image enters Stage I screening for referable diabetic retinopathy at grade 2 or above. It then branches to five parallel lesion segmentation models — microaneurysms, haemorrhages, hard exudates, cotton wool spots and vessels — whose masks rejoin the raw image as eight channels into Stage III, the lesion-guided grader. Stage III emits either an ordinal severity grade from one to four, or an ungradable verdict routed to human review.Fundus imageStage I · screening(referable DR, grade ≥ 2)MAHSMoEHEMultiScaleEXBrightSpotCWSwarm-upVesselsFOV maskStage III · LG-DRG(8 channels: RGB + 5 masks)Severity grade(CORN ordinal, 1–4)Ungradable(routed to human review)
    Fig. 1 — The HierarchiRetina three-stage pipeline: screening, five parallel lesion segmenters, and lesion-guided grading over eight channels.

    FigModelAsFunction

    Book I, Ch. I, Prop. 1permalink
    A model as a parameterised functionAn input enters a box marked f of x semicolon theta and an output leaves it. A second arrow enters the box from below carrying the parameters theta, drawn as a stack of adjustable values. The input arrives from the world; the parameters are chosen by training.the world suppliesthe model returnsxinputf( x ; θ )fixed formŷpredictionθ — the parameters, and the only thing training may change
    Fig. 2 — A model as a function of two arguments. The input comes from the world; the parameters are the only part training is allowed to change.

    FigTensorShape

    Book I, Ch. I, Prop. 2permalink
    The axes of a tensor and what each one meansA block of shape batch by sequence by feature, drawn in three dimensions, beside a legend naming each axis. The same numbers arranged along different axes are different data; the shape is what says which is which.shape (B, T, D)T — rowsD — columnsB — depthB · batchindependent examplesT · sequencepositions in timeD · featurelearned dimensionsTranspose two axes and the numbers survive;the meaning does not.
    Fig. 3 — The axes of a tensor, named. The same numbers arranged along different axes are different data, which is why a shape is a claim about meaning and not only about storage.

    FigParameterSpace

    Book I, Ch. I, Prop. 3permalink
    Training as a path through parameter spaceTwo parameter axes with a sequence of points joined into a path, starting at a randomly initialised point and ending at a settled one. Every point on the path is the same model with different numbers in it.parameter spaceθ₁θ₂θ₀θ*Same architectureat every point.Only the numbersmove.
    Fig. 4 — Training as a path through parameter space. Every point on the path is the same architecture holding different numbers.

    FigLossLandscape

    Book I, Ch. I, Prop. 4permalink
    A loss surface, drawn as contoursNested contour rings around a deep minimum on the left and a shallower one on the right, with a saddle between them. A dashed line runs from a starting point downhill into the nearer basin, not necessarily the better one.loss surface over two parametersgloballocalstartThe loss is the only statement of the objective the optimiser can read. What it omits is not optimised.
    Fig. 5 — A loss surface drawn as contours, with a deep basin and a shallower one. Descent finds the nearer minimum, which is not always the better one.

    FigGeneralisation

    Book I, Ch. I, Prop. 5permalink
    Training error and held-out errorTwo curves against training time. Training error falls steadily. Held-out error falls with it, reaches a minimum, then rises. The vertical distance between the curves after that point is the memorised part.error against training timeepochserrorbest held-outthe gap:what wasmemorisedtrainingheld out
    Fig. 6 — Training error and held-out error against training time. The gap that opens after the held-out minimum is the memorised part.

    FigHyperplane

    Book I, Ch. II, Prop. 1permalink
    A perceptron draws one hyperplane, and w is perpendicular to itA straight boundary divides the plane into two half-spaces, one labelled plus and one minus. The weight vector is drawn as an arrow leaving the boundary at a right angle and pointing into the positive side; the bias slides the boundary along that arrow without turning it.one boundary, two half-spaces⟨w, x⟩ + b = 0w+++−−−score > 0score < 0the only shape a perceptron can drawb slides the line along w · it cannot turn it
    Fig. 7 — The weight vector is the normal to the boundary, so the bias slides the boundary along w and can never turn it.

    FigSignedDistance

    Book I, Ch. II, Prop. 2permalink
    The score divided by the norm of w is a signed distanceA point sits away from the boundary. The perpendicular dropped from it to the boundary has length equal to the score divided by the length of the weight vector; the sign of the score says which side, and the magnitude says how far.one number, two factsboundaryx|⟨w,x⟩+b| / ‖w‖x′sign → the classsize → the distancescores rescale with w; distances do notdoubling w and b doubles every score and moves nothing
    Fig. 8 — One number carries both facts: its sign is the predicted class and its magnitude, once divided by the norm of w, is the distance to the boundary.

    FigPerceptronUpdate

    Book I, Ch. II, Prop. 3permalink
    An update rotates the boundary toward the misclassified pointBefore the update a positive point lies on the negative side. Adding y times x to the weight vector turns the weight vector toward that point, which swings the boundary until the point is on the correct side. Points already classified correctly produce no change at all.beforeafter one updatex, y = +1margin ≤ 0 — a mistakewx, y = +1margin > 0 — correctw + y·xold ww ← w + y·x turns the normal toward the point; the boundary follows
    Fig. 9 — An update turns the weight vector toward the point it got wrong, and the boundary swings with it. A point already correct produces no change at all.

    FigXorSquare

    Book I, Ch. II, Prop. 4permalink
    XOR puts each class on one diagonal of the square, and the diagonals crossThe four corners of the unit square carry XOR's labels. The two positive corners lie on one diagonal and the two negative corners on the other. Any straight line separating the ends of one diagonal must pass between the ends of the other, so no line can separate the classes.XOR on the unit square0,0 → 01,1 → 01,0 → 10,1 → 1b ≤ 0w₂ + b > 0w₁ + b > 0w₁ + w₂ + b ≤ 0add the middle two: b > 0against the first: impossiblefilled = 1 · hollow = 0no straight line separates the filled corners from the hollow ones
    Fig. 10 — XOR puts each class on one diagonal of the square, and the diagonals cross. Any line separating one diagonal must pass between the other.

    FigLossShapes

    Book I, Ch. IV, Prop. 1permalink
    Squared, absolute and Huber losses, with the derivatives each sends backOn the left the three loss curves against the residual: the squared loss rises as a parabola, the absolute loss as a V, and Huber follows the parabola near zero and the V beyond delta. On the right their derivatives: the squared loss's grows without bound, the absolute loss's is a step of plus or minus one, and Huber's rises like the squared loss then flattens.the lossestheir derivatives — what training uses½r² |r|Huber±δunbounded±δthe left panel is what you report · the right panel is what moves the modelHuber's value tracks the parabola; its pull is capped at δ, like the V's
    Fig. 11 — The left panel is the number you report; the right panel is what moves the model. Huber tracks the parabola in value and caps its pull at δ.

    FigLossIsALikelihood

    Book I, Ch. IV, Prop. 2permalink
    Every loss is the negative log-likelihood of some noise modelA table pairing each observation model with the loss it produces once its negative log-likelihood is taken and the constants discarded. Choosing the loss on the right asserts the model on the left, whether or not the assertion was intended.choose a loss and you have chosen a noise modelassumed distributionresulting losswhat it asserts about the dataGaussian, fixed σ²squared errorsymmetric · constant spread · unboundedLaplaceabsolute errorsymmetric · heavy tailsCategoricalcross-entropyone of K, exclusiveBernoullibinary cross-entropyyes or noNegative binomialNB likelihoodcounts · variance grows with mean−log p(y | x, θ), constants discarded — that is the whole derivation
    Fig. 12 — Choosing the loss on the right asserts the distribution on the left, whether or not the assertion was intended.

    FigGradientCancellation

    Book I, Ch. IV, Prop. 3permalink
    Cross-entropy cancels the output derivative; squared loss keeps itTwo chains from a logit to a gradient. Along the top, cross-entropy contributes a factor that is the reciprocal of the sigmoid derivative, so the two cancel and the gradient is simply p minus y. Along the bottom, squared loss contributes no such factor, so the sigmoid derivative survives and shrinks the gradient toward zero exactly where the model is most wrong.cross-entropysquared loss∂L/∂p(p−y) / p(1−p)×dp/dzp(1−p)=p − ythese cancel exactlybounded by 1largest when most wrong∂L/∂pp − y×dp/dzp(1−p)=(p−y)·p(1−p)nothing cancels→ 0 as the model worsensat z = −10 with y = 1: cross-entropy sends −0.99996, squared loss sends −0.0000454
    Fig. 13 — Cross-entropy contributes the reciprocal of the output derivative, so the two cancel. Squared loss contributes nothing, so the derivative survives and shrinks the gradient exactly where the model is most wrong.

    FigSaturation

    Book I, Ch. III, Prop. 2permalink
    The sigmoid derivative peaks at one quarter and collapses either sideA curve of the sigmoid derivative against its input, peaking at one quarter at the origin and falling to almost nothing by an input of four. Beside it, the product of that derivative over increasing depth, reaching the smallest number representable in half precision at twelve layers.σ′ against z0.25 raised to the depthmax = ¼−60+60.250L = 10.25L = 30.016L = 59.8e-4L = 81.5e-5L = 126.0e-8fp16 smallest positive: 6.0e−8at depth 12 the gradient is not small — it is zeroand this is the best case; at |z| = 4 it is depth 5
    Fig. 14 — The sigmoid derivative never exceeds one quarter, and a product of them reaches the smallest half-precision number by depth twelve — in the best case that never occurs.

    FigDeadUnit

    Book I, Ch. III, Prop. 3permalink
    A dead ReLU unit is a fixed point of gradient descentThe whole input domain maps to negative pre-activations, so the unit outputs zero. The backward pass multiplies the downstream error by the ReLU derivative, which is exactly zero there, so no gradient reaches the weights and they never move. The cycle closes on itself.the loop that cannot be brokeninputs in [0,1]²w = (1,1), b = −10z ∈ [−10, −8]always negativea = 0contributes nothing∂L/∂a from downstream× ReLU′(z)= 0 exactly∂L/∂w = 0so w never movesforward above · backward below · the state never changesLeakyReLU replaces the exact zero with 0.01 — slow, but no longer stuck
    Fig. 15 — A dead ReLU unit is a fixed point of gradient descent: the derivative that would move its weights is the same zero that killed it.

    FigSmoothVsKink

    Book I, Ch. III, Prop. 4permalink
    ReLU has a kink where GELU has a curve, and the derivatives differ sharply thereOn the left, ReLU and GELU plotted together: they agree away from the origin and differ in a narrow band around it, where GELU dips below zero. On the right, their derivatives: ReLU's jumps from zero to one with no value at the kink, while GELU's passes smoothly through one half and briefly goes negative.the functionstheir derivativesReLUGELUkink½GELU′(0) = ½ReLU′(0) undefinedGELU′ dips to −0.13 near z = −√2the whole difference lives in a band around the origin — which is where gradients are decided
    Fig. 16 — ReLU and GELU agree away from the origin and differ only in a narrow band around it — which is exactly where the gradient is decided.

    FigNeuron

    Book I, Ch. II, Prop. 1permalink
    A single unit: weighted sum, bias, activationThree inputs each multiplied by a weight and summed, with a bias added, then passed through an activation function to produce one output. The weighted sum and bias are the affine part; only the last box is non-linear.one unitx₁w₁x₂w₂x₃w₃Σb — biasσactivationaaffine: a fixed shape, only the numbers learnedthe only non-linear step
    Fig. 17 — One unit: a weighted sum, a bias, and an activation. Everything before the activation is affine, and affine maps compose into a single affine map.

    FigDepthFolds

    Book I, Ch. II, Prop. 2permalink
    Composition with and without a non-linearityOn the left, two linear layers composed produce a single straight line: the stack collapses to one layer. On the right, the same two layers with a rectifier between them produce a piecewise-linear curve with three segments.linear ∘ linearlinear ∘ relu ∘ linearone line, whatever the depthkinkkinkeach unit contributes a fold
    Fig. 18 — Two linear layers composed give one line whatever the depth. The same two layers with a rectifier between them give a piecewise curve, one fold per unit.

    FigComputationalGraph

    Book I, Ch. II, Prop. 3permalink
    One graph, traversed forwards then backwardsFive nodes in a row: input, a linear step, an activation, a second linear step and the loss. Arrows along the top run left to right carrying values. Arrows underneath run right to left carrying partial derivatives, each one multiplied into the next by the chain rule.forward — valuesbackward — gradientsxu = Wxh = σ(u)ŷ = VhL∂L/∂·∂L/∂·∂L/∂·∂L/∂·Each backward arrow is one local derivative multiplied into what arrived from the right.The cost of the backward pass is the cost of the forward pass, within a small constant.
    Fig. 19 — One graph traversed twice: values forwards, partial derivatives backwards. Backpropagation is the chain rule with the intermediate results kept.

    FigDescentStep

    Book I, Ch. II, Prop. 4permalink
    One step, three learning ratesThe same quadratic bowl drawn three times. With too small a step the parameters creep toward the minimum. With a well-chosen step they reach it. With too large a step they overshoot the minimum and climb the far wall.too smallcreepswell chosenarrivestoo largediverges
    Fig. 20 — The same descent under three learning rates. Too small and it creeps, too large and it climbs the far wall; the gradient supplies the direction, never the distance.

    FigRegularisation

    Book I, Ch. VIII, Prop. 1permalink
    An exact fit and a constrained oneSeven data points fitted by a curve that passes exactly through every one, and by a straighter curve that misses all of them slightly. A held-out point lies far from the exact curve and close to the constrained one.the same seven points, fitted twiceheld outpasses through every pointpasses through none of themRegularisation does notimprove the fit. It decideswhich fit you get whenseveral are available.
    Fig. 21 — Seven points fitted exactly and fitted loosely. Regularisation does not improve the fit — it decides which fit you get when many are available.

    FigNormalisationAxes

    Book I, Ch. VIII, Prop. 2permalink
    Batch and layer normalisation reduce over different axesTwo identical three-by-three matrices have examples in rows and features in columns. Batch normalisation highlights a column: one feature across three examples. Layer normalisation highlights a row: three features within one example. The other columns or rows form their own separate groups. Neither operation changes the tensor shape.batch normalisationfeatures →147258369examplesone feature, across examplesOther examples affect this output.Inference normally freezes the statistics.layer normalisationfeatures →147258369examplesone example, across featuresOther examples do not affect it.The same rule at training and inference.
    Fig. 22 — The same matrix, different reduction groups. Batch normalisation shares statistics down a feature column; layer normalisation shares them across one example. Highlighted entries form one group; the remaining rows or columns form their own groups.

    FigTokenBoundary

    Book I, Ch. III, Prop. 1permalink
    One string under three tokenisationsThe string "unhappiness" segmented three ways: as subwords un / happi / ness, as a single word token, and as ten individual characters. Each segmentation fixes which distinctions the model is able to represent at its input.unhappinesssubwordunhappinesswordunhappinesscharunhappiness
    Fig. 23 — One string under three tokenisations. The segmentation chosen at the input fixes which distinctions the model is able to represent at all.

    FigBPEMerge

    Book I, Ch. III, Prop. 2permalink
    Byte-pair encoding, three merges deepThe word lowest, first split into single characters, then merged three times. Each row shows the vocabulary after one merge of the most frequent adjacent pair, ending with the two pieces low and est.most frequent adjacent pair, mergedstartlowestmerge 1lowestmerge 2lowestmerge 3lowestvocabulary: 6+ "st" → 7+ "lo" → 8+ "low", "est" → 10The merge list is the tokeniser. Nothing in it is learned, and it is fixed before training begins.
    Fig. 24 — Byte-pair encoding, three merges deep. The merge list is the tokeniser; it is fixed by frequency in a corpus before training begins, and nothing in it is learned by the model.

    FigEmbeddingSpace

    Book I, Ch. III, Prop. 3permalink
    The embedding table and the space it inducesOn the left a table mapping token ids to rows of numbers, with an arrow picking out one row. On the right, four of those rows drawn as points, with the offset between king and queen matching the offset between man and woman.a lookupand what the rows come to mean4711d numbers, learned 1204d numbers, learned 0987d numbers, learned One row per vocabulary entry. Retrieval is indexing, not computation.kingqueenmanwomanThe geometry is a consequence oftraining, not of the lookup.
    Fig. 25 — The embedding table and the space it induces. Retrieval is indexing; the geometry among the retrieved rows is a consequence of training.

    FigContextualisation

    Book I, Ch. III, Prop. 4permalink
    One embedding, two contextsTwo sentences that share the word bank. In both, the token enters the model as the identical vector from the embedding table. After a layer of attention the two copies have moved apart, because each has read a different context.at the input — identicaltheriverbankfloodedthecentralbankraisedone layer of attentionafter — separatedbank (river)bank (finance)context supplied the difference
    Fig. 26 — One token in two sentences. It enters as the same vector in both, and only separates once a layer of attention has let it read its neighbours.

    FigRecurrentState

    Book I, Ch. IV, Prop. 1permalink
    A recurrence unrolled over five stepsFive inputs entering five copies of the same cell, each passing a fixed-width hidden state to the next. Everything read so far has to fit through that one state, whose width does not grow with the sequence.the same cell, five timesx₁fh1y1x₂fh2y2x₃fh3y3x₄fh4y4x₅fy5The state has a fixed width. Step 5 reaches step 1 only through everything between it and step 1.
    Fig. 27 — A recurrence unrolled. Every step passes a fixed-width state to the next, so the path between two distant positions is as long as the distance between them.

    FigVanishingGradient

    Book I, Ch. IV, Prop. 2permalink
    A gradient carried back through many stepsThree curves on a logarithmic scale showing a repeated multiplication over sixteen steps. A factor below one falls away to nothing, a factor above one grows without bound, and only a factor of exactly one holds steady.magnitude after k steps, log scalesteps back0.7 — vanishes1.0 — steady1.3 — explodesNot specific to recurrence: this is what happens to any long product of numbers that are not exactly one.
    Fig. 28 — A repeated multiplication over sixteen steps. Nothing here is peculiar to recurrence: it is what happens to any long product of numbers that are not exactly one.

    FigGating

    Book I, Ch. IV, Prop. 3permalink
    A gated cell and its uninterrupted carry lineA horizontal line across the top carries the cell state from left to right, crossed only by one multiplication and one addition. Below it, a forget gate and an input gate produce the numbers that do the multiplying and adding.the carry linec(t−1)c(t)×+forget gateσ → [0, 1]input gatewhat to writeh(t−1), x(t) — both gates read the same two thingsWith the forget gate near one the carry line is near the identity, and a gradient crossing it barely decays.
    Fig. 29 — A gated cell. The carry line crosses only a multiplication and an addition, so with the forget gate near one the gradient has a route home that no weight matrix attenuates.

    FigDilation

    Book I, Ch. IV, Prop. 4permalink
    Receptive field growth under stacked dilated convolutionsFour rows of positions. Each layer reads three positions from the row below at a spacing that doubles, so the window seen by the top unit widens from three positions to fifteen across three layers.inputlayer 1dil 1 · sees 3layer 2dil 2 · sees 7layer 3dil 4 · sees 15Range grows exponentially with depth, and every position is computed at once.
    Fig. 30 — Receptive field growth under stacked dilated convolutions. Range grows exponentially with depth and every position is computed at once — attention buys the same range at depth one.

    FigSequenceLineage

    Book I, Ch. IV, Prop. 5permalink
    Four ways to read a sequence, by path length and parallelismA table of four architectures. Recurrent models connect distant positions through a path that grows with distance and cannot be parallelised. Convolution shortens the path logarithmically and parallelises. Attention connects any two positions in one step, in parallel, at quadratic cost.architecturepath between two positionsover the sequenceRNNO(n)sequentialLSTM / GRUO(n)sequentialCNNO(log n)parallelAttentionO(1)parallelAttention is not a better idea than recurrence in the abstract. It is the trade that pays when thehardware is parallel and the sequences are short enough to afford n².
    Fig. 31 — Four ways to read a sequence, placed by path length and parallelism. Attention is not a better idea in the abstract; it is the trade that pays on parallel hardware.

    FigAttentionWeights

    Book I, Ch. V, Prop. 1permalink
    Attention as a weighted average over valuesA single query is compared against four keys. The resulting softmax weights — 0.06, 0.61, 0.09 and 0.24 — are shown as horizontal bars, and the output is the sum of the value vectors scaled by those weights. The weights come from content, not from position.query qsoftmax(q · kᵢ / √d)k₁0.06k₂0.61k₃0.09k₄0.24Σ wᵢ vᵢweights are content-addressed — nothing here depends on i
    Fig. 32 — Attention as a weighted average over values, with weights computed by comparing a query against every key. Nothing in the computation depends on position.

    FigQKV

    Book I, Ch. V, Prop. 2permalink
    Query, key and value as three projections of one vectorA single token vector on the left fans out into three learned projections. The query asks what the token is looking for, the key advertises what it offers, and the value is what gets passed on once a match is found.xone tokenWQQquerywhat am I looking for?WKKkeywhat do I advertise?WVVvaluewhat do I hand over?Q and K decide the weights. V is the only one that reaches the output.
    Fig. 33 — Query, key and value as three learned projections of one vector: what a token is looking for, what it advertises, and what it hands over once matched.

    FigSoftmaxTemperature

    Book I, Ch. V, Prop. 3permalink
    The same attention logits, scaled and unscaledTwo bar charts of attention weights over eight keys. Scaled, the weights are spread across several keys. Unscaled, almost all the weight lands on one key and the remaining bars are invisible, which is where the gradient dies.scaled by 1/√dk1k2k3k4k5k6k7k8max weight 27%unscaled — logits grow with dk1k2k3k4k5k6k7k8max weight 80%For unit-variance q and k the dot product has variance d, so the logits grow with the head width.Dividing by √dₖ restores it, whatever the width.
    Fig. 34 — The same attention logits, scaled and unscaled. Without the divisor the distribution collapses onto one key and the gradient through the softmax goes flat.

    FigSelfVsCross

    Book I, Ch. V, Prop. 4permalink
    Self-attention and cross-attention are one operationTwo identical attention blocks. In the first, queries, keys and values all come from the same sequence. In the second, the queries come from one sequence and the keys and values from another. The block itself is unchanged.self-attentionthe sequencethe same sequenceQK, Vattentionone vector per querycross-attentionthe decoderthe encoderQK, Vattentionone vector per query
    Fig. 35 — Self-attention and cross-attention are one operation under two wirings. Only the source of the keys and values changes.

    FigAttentionCost

    Book I, Ch. V, Prop. 5permalink
    The attention score matrix at three sequence lengthsThree square grids of scores, for sequences of four, eight and sixteen tokens. Each doubling of the sequence length quadruples the number of scores, because every token is compared against every other.one score per pair of positionsn = 416 scoresn = 864 scoresn = 16256 scoresPath length is constant — any token reaches any other in one step — and the square is what it costs.
    Fig. 36 — The score matrix at three sequence lengths. Constant path length between any two positions is bought with an area that grows as the square of the sequence.

    FigPermutation

    Book I, Ch. VI, Prop. 1permalink
    Permutation equivariance, and the repairIn the upper panel, self-attention alone maps a shuffled input to an identically shuffled output: reordering the tokens carries no information. In the lower panel the same tokens carry positional encodings, so a reordering produces a genuinely different output.attention aloneattention + positionthecatsatsatthecatthecatsatsatthecat+p1+p2+p3+p1+p2+p3the′cat′sat′sat′the′cat′the′cat′sat′cat′sat′the′same output, merely reordered — the model cannot tell the two inputs apartdifferent output — order is now information
    Fig. 37 — Permutation equivariance and its repair. Without positional encoding a reordered input yields a merely reordered output; with it, order becomes information.

    FigSinusoid

    Book I, Ch. VI, Prop. 2permalink
    Sinusoidal position encoding as a bank of clocksThree sine waves of geometrically increasing wavelength plotted against position. Together they give every position a distinct signature: the fast wave distinguishes adjacent positions, the slow one distinguishes distant regions of the sequence.one dimension per wavelengthfast — separates neighboursmiddleslow — separates regionsone position, read off every clock at once
    Fig. 38 — Sinusoidal encoding as a bank of clocks. Fast dimensions separate neighbouring positions, slow ones separate distant regions, and together they give each position a signature.

    FigRoPE

    Book I, Ch. VI, Prop. 3permalink
    Rotary position embeddingA circle with three vectors drawn at angles proportional to their positions in the sequence. Because both the query and the key are rotated, the angle between any two of them depends only on how far apart the positions are.position becomes an anglem = 2m = 5m = 9⟨ R(mθ)q , R(nθ)k ⟩= ⟨ q , R((n−m)θ)k ⟩The absolute indices cancel. What survivesthe inner product is n − m, the distance.Nothing is added to the embedding, so theresidual stream is left alone; only q and kare turned, inside each head.
    Fig. 39 — Rotary encoding turns position into an angle. Both the query and the key are rotated, so the absolute indices cancel in the inner product and only the distance survives.

    FigExtrapolation

    Book I, Ch. VI, Prop. 4permalink
    Beyond the training lengthA position axis with a marked training length. A learned absolute encoding is drawn as a row of slots that simply stops at that boundary. A sinusoidal or rotary encoding continues past it, but into a region the model has never been trained on.learned absoluteno row exists — the table ends heresinusoidal / rotarydefined, but never trained at these distancestraining lengthA long-context claim is a claim about what was trained, not about what the encoding can be evaluated at.
    Fig. 40 — Past the training length a learned table has no row at all, and a periodic scheme has a value but no experience. A long-context claim is a claim about training, not about the encoding.

    FigTransformerBlock

    Book I, Ch. VII, Prop. 1permalink
    One transformer blockA block with two sublayers. The first normalises, applies multi-head attention across positions and adds the result back to the input. The second normalises, applies a position-wise feed-forward network and adds again.mix across positionsthen compute within onexoutnormmulti-head attention+the residual path — the block edits, it does not replacenormfeed-forward (per token)+no token sees another here
    Fig. 41 — One transformer block: attention moves information between positions, the feed-forward network does the work within one, and both are wrapped in a residual addition.

    FigMultiHead

    Book I, Ch. VII, Prop. 2permalink
    One attention budget, divided into four headsA vector of width d split into four slices of width d over four. Each slice runs its own attention and attends to a different relation. The four outputs are concatenated back to width d and mixed by an output projection.one vector, width dhead 1d / 4previous tokenhead 2d / 4the subjecthead 3d / 4matching quotehead 4d / 4nothing usefulconcatenate, then W_OHeads are cheap because they are slices. What they buy is several relations at once;what they cost is resolution inside each.
    Fig. 42 — One attention budget divided into four heads. Heads are slices rather than copies, so several relations are attended to at once at the cost of resolution inside each.

    FigResidualStream

    Book I, Ch. VII, Prop. 3permalink
    The residual streamA single horizontal line running the width of the figure, from the embedding to the output head. Five blocks sit beside it; each reads from the line and adds its result back onto it rather than replacing it.the residual streamembedunembedblock 1readwriteblock 2readwriteblock 3readwriteblock 4readwriteblock 5readwriteBecause every block adds, the gradient reaches layer one along a path of additions, not a product.And because every block writes into the same space, a vector taken partway up the stack can be readwith the output head — which is what makes a transformer legible from the inside at all.
    Fig. 43 — The residual stream. Blocks read from it and add back into it rather than replacing it, which is both why gradients reach layer one and why intermediate states can be read.

    FigPrePostNorm

    Book I, Ch. VII, Prop. 4permalink
    Post-norm and pre-normTwo arrangements of the same sublayer and normalisation. In post-norm the residual addition is followed by a normalisation, so the shortcut passes through it. In pre-norm the normalisation happens before the sublayer and the shortcut runs clean from input to output.post-normthe shortcut passes through a normsublayer+normshortcutneeds a warm-up; deep stacks are fragilepre-normthe shortcut is untouchednormsublayer+shortcuttrains from a flat start; stacks deeply
    Fig. 44 — Post-norm puts a normalisation on the shortcut; pre-norm leaves the shortcut clean. The same two components in two orders, with different training behaviour.

    FigCausalMask

    Book I, Ch. VII, Prop. 5permalink
    The causal maskA square grid of attention scores for seven tokens. The lower triangle, including the diagonal, is open: each token may attend to itself and to everything before it. The upper triangle is closed, because those positions lie in the future.query row attends to key columnthe−∞−∞−∞−∞−∞−∞cat−∞−∞−∞−∞−∞sat−∞−∞−∞−∞on−∞−∞−∞the−∞−∞warm−∞matthecatsatonthewarmmatSet before the softmax, so themasked entries receive exactlyzero weight rather than a small one.This is what lets every position betrained at once on the same sequence:n prediction problems, one forward pass.Remove the triangle and the same weights become an encoder. Nothing else about the block changes.
    Fig. 45 — The causal mask. Setting the future to minus infinity before the softmax is what lets every position in a sequence be trained as a separate prediction in one pass.

    FigEncoderDecoder

    Book I, Ch. VII, Prop. 6permalink
    Encoder, decoder, and the two joinedThree arrangements of the same block. An encoder attends over the whole input. A decoder attends only over what precedes each position. An encoder-decoder runs both and joins them with a cross-attention layer.encoder-onlyencodersees: the whole inputused for: classify, embed, segmentdecoder-onlydecodersees: the past onlyused for: generateencoder–decoderencoderdecodercross-attentionsees: both, joinedused for: translate, captionSame block, same maths. The masks and the wiring are the whole of the difference.
    Fig. 46 — Encoder, decoder, and the two joined by cross-attention. Same block, same arithmetic — the masks and the wiring are the whole of the difference.

    FigBatchNoise

    Book I, Ch. VIII, Prop. 1permalink
    The batch gradient as an estimateTwo panels, each showing the true gradient direction as a solid arrow and several batch estimates around it. With a small batch the estimates scatter widely; with a large batch they cluster near the true direction.batch of 8noisy direction, many cheap stepsbatch of 512clean direction, few costly stepsNoise in the estimate is not purely a defect: it is also what lets a small batch escape a shallowbasin that a large one would settle into.
    Fig. 47 — The batch gradient as an estimate of the true one. A small batch scatters widely and steps often; a large batch points true and steps rarely.

    FigSchedule

    Book I, Ch. VIII, Prop. 2permalink
    Warm-up followed by cosine decayLearning rate against training step. The rate rises linearly from zero over a short warm-up, reaches a peak, then decays along a cosine curve to near zero by the end of training.learning rate against stepwarm-updecaypeak ratethe rise exists because the first steps are takenby an optimiser with no statistics yetthe fall exists because a large step near the endthrows away what the previous ones foundChanging the step count changes the whole curve. A schedule tuned for one budget is not valid for another.
    Fig. 48 — Warm-up then decay. The rise exists because the optimiser has no statistics yet; the fall exists because a large step near the end discards what the earlier ones found.

    FigClipping

    Book I, Ch. VIII, Prop. 3permalink
    Weight decay and gradient clipping guard different thingsA bar chart of gradient norms over fourteen steps, mostly small with two enormous spikes. A horizontal clipping threshold cuts only the two spikes. Beneath, a separate note shows weight decay acting on every step regardless.gradient norm per stepclip thresholdClipping is a rare intervention: on almost every step it does nothing at all.Weight decay is the opposite: a small constant pull toward zero, applied on every step to everyparameter, whatever the gradient is doing. One bounds the step, the other bounds where they land.
    Fig. 49 — Gradient norms over fourteen steps with a clipping threshold. Clipping touches only the rare spike; weight decay pulls on every parameter every step. They guard different failures.

    FigPrecision

    Book I, Ch. VIII, Prop. 4permalink
    Representable range in three float formatsThree horizontal bands on a logarithmic magnitude axis. Half precision covers a narrow band; brain float and single precision cover a wide one. A marked region of small gradient magnitudes falls below the half-precision floor and rounds to zero.magnitude, log scalefp32fp16bf16where gradients livelate in trainingLoss scaling multiplies the loss by a largeconstant so the gradients land inside the band,then divides it back out before the update.
    Fig. 50 — Representable range in three float formats. Half precision is narrow, and late-training gradients fall through its floor — which is what loss scaling exists to prevent.

    FigSplitLeak

    Book I, Ch. VIII, Prop. 5permalink
    Two ways to split the same six imagesSix images from three patients, two each. Split by image, every patient appears on both sides of the split and the held-out score is inflated. Split by patient, no patient appears on both sides and the score is honest.split by imagetrainp1p2p3held outp1p2p3patients 1, 2 and 3 all appear on both sidesmeasures recall of patients already seensplit by patienttrainp1p1p2p2held outp3p3no patient appears on both sidesmeasures what happens on a new patientThe unit of the split has to be the unit the claim is about — patient, site, scanner, study. Get itwrong and every other measure of rigour is decoration: the number was decided before training began.
    Fig. 51 — The same six images split two ways. Splitting by image puts every patient on both sides; splitting by patient is the only one of the two that measures what the claim is about.

    FigTransformerMap

    Book I — closingpermalink
    The whole of Book I on one plateA vertical map from raw text to output logits. Text is tokenised, embedded and given position, then passed through a stack of transformer blocks, each containing normalisation, attention, a residual addition, a second normalisation and a feed-forward network. The final state is unembedded to logits. Each stage is labelled with the chapter that covers it.from characters to logitstexttokeniserCh. IIItoken idsCh. IIIembedding + positionCh. III, VI× N blocksnormCh. VIImulti-head attentionCh. V, VIIresidual addCh. VIInormCh. VIIfeed-forwardCh. II, VIIresidual addCh. VIIstreamunembed → logitsCh. I, VIIIEverything in Book I ison this plate. If a stagehere is not one you couldexplain to somebody else,that is the chapter toread again.
    Fig. 52 — The whole of Book I on one plate, from characters to logits, with the chapter covering each stage named beside it.

    FigGramAnchoring

    Marginalia II — DINOv2 and DINOv3permalink
    Patch feature quality over a long training runDense task quality plotted against training steps. Early on it rises and holds. Left alone it then declines as the image-level objective pulls the representation toward abstraction. A later phase that pins the patch-to-patch Gram matrix to an earlier teacher checkpoint recovers it.dense task quality against training stepsanchoring beginsanchored — +2 mIoUleft alone — features rotScaling to 7B only paid off because they fixed this first.
    Fig. 53 — Dense features rot as a large ViT trains, because the image-level objective wants abstraction and wins. Anchoring the patch-to-patch Gram matrix to an earlier checkpoint drags them back.

    FigFcmaeGrn

    Marginalia III — ConvNeXt V2permalink
    Masking a convolution, and the collapse it causesOn the left, a masked grid with a kernel window straddling the boundary: an ordinary convolution reads across the hole. In the middle, sparse convolution treats visible patches as sparse data. On the right, channel responses collapse toward each other until global response normalisation forces them to compete.the problemFCMAEGRNa kernel spans the holeand leaks the answervisible patches only,as sparse datacollapsedcompetingchannels forced apartby normalisationNeither half works alone. Bolting MAE onto plain ConvNeXt made things worse.
    Fig. 54 — A kernel slides across a masked hole and leaks the answer; sparse convolution stops it, and global response normalisation stops the channels collapsing. Neither half works alone.

    FigConvHandback

    Marginalia IV — Swin UNETR V2permalink
    One encoder stage, before and afterThe V1 encoder stage is two Swin transformer blocks followed by patch merging. V2 prepends a residual convolution to every stage. That single block is the entire contribution of the paper.Swin UNETRSwin UNETR V2Swin blockSwin blockpatch mergingresidual convSwin blockSwin blockpatch mergingin MONAI this is use_v2=TrueSelf-attention has no stronginductive bias, so the model isharder to train and hungrierfor data — their own abstract.
    Fig. 55 — One residual convolution at the head of every encoder stage. That single block is the entire contribution of the paper.

    FigLinearProbe

    Marginalia V — logistic regressionpermalink
    What a linear probe actually measuresA frozen encoder feeds features into a single linear layer and a softmax. The encoder is the thing under test; the classifier on top is logistic regression, and it is the instrument doing the measuring.the thing being testedthe instrumentencoder, frozenDINOv2 · ConvNeXt V2 · SigLIPno gradient reaches herefeaturesWx + b→ softmaxone weight per featurepConvex loss. One optimum.The same answer every run,which is why it can be a ruler.Every self-supervised paper reports this number. The model doing it is the one nobody names.
    Fig. 56 — What a linear probe measures, and what does the measuring. The encoder is frozen; the classifier on top is logistic regression, and its convexity is why it can serve as a ruler at all.

    FigFingerprint

    Marginalia VI — nnU-Netpermalink
    Dataset fingerprint to pipelineFive measured properties of a dataset enter a set of heuristic rules, which resolve them jointly under a GPU memory budget into five pipeline decisions. The architecture is not among the things being chosen.fingerprintpipelinevoxel spacingsimage sizesintensity distributionmodalityclass ratiosheuristic rulesunder a memory budgettarget spacingresamplingnormalisationpatch and batch sizedepth and poolingFixed regardless: Dice plus cross-entropy, SGD at 0.99, 1000 epochs, the same augmentations every time.The original paper used no residual connections, no attention, no squeeze-and-excitation. That was the claim.
    Fig. 57 — A dataset fingerprint resolved by heuristic rules into a whole pipeline under a memory budget. The architecture is not among the things being chosen.

    FigScalingFlat

    Marginalia VII — EVA-CLIPpermalink
    Zero-shot average across 27 benchmarksThree bars on a truncated axis. Going from eight billion parameters to eighteen billion gains seven tenths of a point. Raising the input resolution on the smaller model, with no extra parameters at all, gains six.zero-shot average, 27 benchmarks79.4%EVA-CLIP-8B224px80%EVA-CLIP-8B448px, same params80.7%EVA-CLIP-18B224px, 2.2x params78.5Axis truncated. The data stayed fixed the whole way — scaling parameters still works, it just works slowly,and the lever everyone else is pulling is the data.
    Fig. 58 — Doubling the parameters on a fixed dataset bought seven tenths of a point. Raising the input resolution, with no extra parameters at all, bought six.

    FigQwenLadder

    Marginalia VIII — Qwenpermalink
    The open ladder and the closed rungFive rungs of a model ladder, each wider than the last. The four lower ones ship as open weights. The top rung stayed API-only for most of the family's life, and the gap between the two is the measure of the commitment.ship the whole ladder at oncelaptopopen weightsworkstationopen weightsserveropen weightsclusteropen weightsMax classAPI only, until Aug 2026Watch what they hold back, not what they release. That gap has moved in both directions.
    Fig. 59 — The whole ladder shipped at once, with the top rung held back. The gap between the open rungs and the closed flagship is the honest measure of the commitment.

    FigOpenVsApi

    Marginalia IX — Qwen3.8permalink
    The two checkpoints in one releaseA comparison of the 2.4 trillion parameter flagship against the 27 billion parameter model released beside it. The headline belongs to the first; the second is the one most people can actually run.Qwen3.8-2.4T-A95BQwen3.8-27Bparameters2.4T, 95B active27.8B denseinputtext onlymultimodalthinkingalways onswitchablelicencebespoke, revenue clausesApache 2.0you can host itnoyesThe open checkpoint is not the API model, and the licence is not Apache whatever the headlines said.
    Fig. 60 — The two checkpoints in one release. The headline belongs to the 2.4 trillion; the 27B is the one that fits on hardware you can buy, and it is the Apache one.

    FigPresenceToken

    Marginalia X — SAM 3permalink
    The presence tokenObject queries used to answer two questions at once: whether the concept is present and where it is. SAM 3 gives presence its own global token, leaves localisation to the queries, and multiplies the two scores.one prompt, two questions"striped cat"presence tokenis the concept here at all?object querieswhere, for each instance×scoreTwo questions that interfered when one set of queries answered both. cgF1 55.7 against a field near 24.
    Fig. 61 — Whether the concept is present and where it is are different questions that used to interfere. Giving presence its own token and multiplying the scores is most of why the numbers moved.

    FigKvCache

    Marginalia XV — DeepSeek V4permalink
    Cost per token at a one-million-token contextTwo bars per model. Compared with the previous generation, V4 needs about twenty-seven percent of the per-token inference operations and ten percent of the key-value cache at a one-million-token context.at 1M context, relative to V3.2V3.2FLOPs100%KV cache100%V4FLOPs27%KV cache10%Compressed Sparse Attention interleaved with Heavily Compressed Attention. Everyone quoted the 1.6 trillion.The number that matters is the 10%.
    Fig. 62 — At a one-million-token context, roughly a quarter of the per-token operations and a tenth of the key-value cache. Everyone quoted the parameter count.

    FigPricePerTask

    Marginalia XVI — Kimi K3permalink
    Rate card against cost per taskThree models compared twice: by their published price per million input tokens, and by measured cost to complete a task. The ordering is not the same, because a model whose answers run token-efficient pays less per job than its rate card suggests.price per million inmeasured cost per taskKimi K3$3.00$0.94GPT-5.6 Sol$2.20$1.04Opus 4.8$4.50$1.80Three times the rate of K2.6 and a lower bill than either rival. The question is not whether K3 is better.It is whether your workload notices.
    Fig. 63 — Rate card against measured cost per task. A higher price per million tokens and a lower bill per job are not contradictory when the answers run token-efficient.

    FigTotalVsActive

    Marginalia XVII — Mixture of expertspermalink
    Total parameters against active parametersA router scores one token against many expert feed-forward layers and runs only the top few. Every expert must be stored in memory; only the fired ones do arithmetic. Compute tracks the small number, memory the large one.one tokentokenroutertop-kevery expert stored — memorytwo fired — computeA parameter count isno longer a capabilityclaim. It is a hardwarerequirement.
    Fig. 64 — Every expert is stored; only the routed few compute. Memory tracks the big number and arithmetic tracks the small one, which is the whole reason anyone does this.

    FigContextClaim

    Marginalia XVIII — Llama 4permalink
    Long-context comprehension at 120K tokensAccuracy on long-context comprehension at a hundred and twenty thousand tokens. Scout, which advertises a ten million token window, scores fifteen point six percent, well under a competitor at the same length.comprehension at 120K tokensGemini 2.5 Pro90.6%Maverick28.1%Scout15.6%Scout advertises 10,000,000 tokens. Needle-in-a-haystack retrieval across that window is genuinely strong,and retrieval is not comprehension.
    Fig. 65 — Comprehension at 120K tokens against an advertised ten-million-token window. Retrieval across a context is not comprehension of it.

    FigFlopsVsLatency

    Marginalia XX — EfficientNetpermalink
    Fewer operations, more time, same accuracyTwo models that both reach 84.0% on ImageNet. One uses 1.8 times fewer floating-point operations and runs 2.7 times slower on the same hardware, because depthwise convolutions are bound by memory bandwidth rather than arithmetic.both land at 84.0% on ImageNetEfficientNet-B6FLOPsarithmetictimewall clockResNet-RS-350FLOPsarithmetictimewall clockAccelerators are bound by memory bandwidth, not arithmetic. "Efficient" meant FLOP-efficient, and a decadeof practitioners read it as fast.
    Fig. 66 — Two models at the same accuracy: one with 1.8x fewer operations and 2.7x more wall-clock time. Accelerators are bound by memory bandwidth, not arithmetic.

    FigContextResolution

    Marginalia XXII — AlphaGenomepermalink
    Context length against output resolutionTwo axes: sequence context and output resolution. Earlier models occupy one corner or the other — fine resolution over a short window, or a long window with binned outputs. AlphaGenome occupies the corner that was assumed unreachable.contextresolution ↑ · context →SpliceAI, BPNet10kb, base resolutionEnformer, Borzoi200–500kb, 32–128bp binsAlphaGenome1Mb, base resolutionthe corner nobody hadThe constraint was computational rather than conceptual. Eleven output modalities at once, from one sequence.
    Fig. 67 — Context against output resolution. Every earlier model sits on one edge or the other; the claim is that the corner between them was an engineering limit rather than a law.

    FigAffinityGap

    Marginalia XXIII — Boltz-2permalink
    Structure, affinity, and the speed between themCo-folding answers whether two molecules fit and runs fast. Free energy perturbation answers how tightly they hold and runs slowly enough to be a scheduling decision. Boltz-2 reaches comparable correlation to the second at the speed of the first.what each method answersco-foldingcan they fitfastFEPhow tightlyslowBoltz-2how tightlyfastPearson 0.62, >1000x fasterApproaching FEP is not matching FEP. A correlation of 0.62 supports ranking a library and choosing what tosynthesise. It does not support telling a chemist that compound 47 binds at 12 nanomolar.Knowing which decisions a number can carry is most of the skill in using it.
    Fig. 68 — Co-folding says whether two molecules fit, free energy perturbation says how tightly and runs slowly. Reaching the second at the speed of the first is what changes screening.

    FigDataCrossover

    Marginalia XXIV — ViTpermalink
    Accuracy against pretraining set sizeTwo curves against the number of pretraining images on a log scale. The convolutional network is ahead on small datasets. The transformer starts lower, rises faster, and overtakes somewhere around a hundred million images. The crossover is the paper's finding.accuracy against pretraining images1M10M100M1BcrossoverConvNet — prior includedViT — prior learnedAbove the threshold theconstraint costs morethan it returns.The claim was conditional, with a threshold in it. What propagated was "transformers beat CNNs".
    Fig. 69 — Accuracy against pretraining set size. The paper's finding is a crossover with a threshold in it, somewhere near a hundred million images — not a verdict.

    FigCountLikelihood

    Marginalia XXV — scVIpermalink
    Where scVI departs from a textbook autoencoderThree rows comparing an ordinary variational autoencoder with scVI: the likelihood is a negative binomial rather than a Gaussian, batch identity conditions both encoder and decoder rather than being corrected afterwards, and library size gets its own latent instead of contaminating the cell state.textbook VAEscVIcount likelihoodGaussian, squared errornegative binomialbatchcorrected afterwardsconditioned on, in both halvessequencing depthleft in the latentits own scaling factorAny method that ignores these produces a beautiful embedding of your experimental logistics.Three independent evaluations, different teams, different data, put this eight-year-old VAE ahead.
    Fig. 70 — Three departures from a textbook autoencoder, each one a fact about the assay: a count likelihood, batch as a conditioning variable, and library size given its own latent.

    FigTwoBenchmarks

    Marginalia XXVI — STATEpermalink
    Two evaluations, opposite conclusionsArc reports that STATE is the first model in its domain to consistently beat simple linear baselines. VCBench reports that pre-registered baselines match or exceed every foundation model tested on four of five dimensions. Four ordinary differences account for the gap.Arc, June 2025first to consistently beatsimple linear baselinesVCBench, June 2026baselines match or exceed allfive models on four of fivebothin printdifferent baselinesdifferent splitsdifferent metricsa year apartNeither claim is dishonest. Working out how they can both be true is more useful than picking a side.If a benchmark cannot tell interaction structure from main effects, a good score on it demonstrates little.
    Fig. 71 — Two credible evaluations pointing opposite ways, and the four ordinary differences that account for it. Neither claim is dishonest.

    FigStripedHyena

    Marginalia XXVII — Evo 2permalink
    Operators by the range they coverFour operator types striped through the model, each covering a different range: short and medium convolutions for local motifs, long implicit convolutions for domain-scale structure, and self-attention reserved for the sparse long-range relationships that need it.one megabase, single-nucleotide resolutionshort explicit convmotifs, splice sitesmedium regularized convlocal regulatory syntaxlong implicit convdomain-scale structureself-attentionenhancer to promoter, sparseAttention costs the square of the length, so you do not pay it across the whole sequence.
    Fig. 72 — Operators assigned by the range they cover, so attention is spent only on the sparse long-range relationships that need it. A megabase becomes affordable.

    FigImageTensor

    Book II, Ch. I, Prop. 1permalink
    An image as a stack of channel planesThree planes of the same height and width, one per colour channel, drawn separated so the axes are visible. Together they are a single tensor of shape height by width by channels; each entry is one number.shape (H, W, C)RGBH — rows of pixelsW — columns of pixelsC — channels, three hereone number per (row, column, channel)A greyscale scan has C = 1.A CT volume adds a depth axis.Nothing else about the model changes.The numbers are only numbers. The shape is the claim about what they are.
    Fig. 73 — An image as a stack of channel planes. The numbers are only numbers; the shape is the claim about what they are.

    FigShuffleDestroys

    Book II, Ch. I, Prop. 2permalink
    The same pixels, reorderedTwo grids holding an identical multiset of pixel values. On the left they are in their original positions and a shape is visible. On the right the same values are permuted and nothing is. Every summary statistic of the two grids is identical.the same values, arranged and shuffledpermutemean, variance, histogram: identicalmean, variance, histogram: identicalA fully connected layer cannot tell these apart at initialisation. The arrangement is the information.
    Fig. 74 — The same pixel values in a different arrangement. Every summary statistic is identical and one of the two is a picture.

    FigNormalisation

    Book II, Ch. I, Prop. 3permalink
    Intensity distribution before and after normalisationA histogram of pixel intensities shifted off centre and narrow, and the same histogram after subtracting the mean and dividing by the standard deviation. The statistics used must come from the training set, and the same ones must be applied at inference.raw intensitiesafter (x − μ) / σ00μ and σ come from the training set, and the same two numbers must be applied at inference. Recomputing themper batch at test time leaks information between test examples, and the leak is silent.A model trained on one scanner's intensity distribution has been told what to expect.Another scanner is a different promise.
    Fig. 75 — Intensities before and after normalisation. The statistics must come from the training set and be applied unchanged at inference.

    FigKernelSlide

    Book II, Ch. II, Prop. 1permalink
    One kernel applied at every positionA six by six input grid with a three by three kernel window that moves across it, and a four by four output grid beside it. The same nine weights are used at every position, which is what weight sharing means and why the parameter count does not grow with the image.inputfeature mapsame 9 weightsSix by six in, four by four out: a 3×3 kernel with no padding loses one position at each edge.Nine weights and a bias, whatever the size of the image. A dense layer here would need over a thousand.That reuse is the locality prior, welded into the architecture rather than learned.
    Fig. 76 — One kernel applied at every position. Nine weights and a bias, whatever the size of the image — the reuse is the locality prior.

    FigPaddingStride

    Book II, Ch. II, Prop. 2permalink
    Output size under three settingsA seven by seven input under three configurations of a three by three kernel. Without padding the output shrinks to five. With padding of one it keeps its size. With a stride of two it halves.7×7 input, 3×3 kernelno padding, stride 1output5 × 5padding 1, stride 1output7 × 7padding 1, stride 2output4 × 4out = ⌊(in + 2·pad − k) / stride⌋ + 1. Nothing else decides it, and it is the commonest shape bug.
    Fig. 77 — A 7×7 input under three settings. Padding and stride decide the output size and nothing else does.

    FigReceptiveGrowth

    Book II, Ch. II, Prop. 3permalink
    Receptive field growth with depthFour rows of positions. One unit at the top depends on three positions in the layer below, seven two layers down, and fifteen three layers down. A 3 by 3 kernel widens the window by two positions per layer, so range is bought with depth.what one unit at the top can seeinputlayer 1sees 5layer 2sees 9layer 3sees 13A 3×3 kernel adds two positions per layer, so the window grows linearly with depth — which is why deep stacksand dilation both exist. Attention reaches the whole image at layer one, and pays a square for it.
    Fig. 78 — What one unit at the top can see, opening by two positions a layer. Range is bought with depth.

    FigPooling

    Book II, Ch. II, Prop. 4permalink
    Two activations, one pooled resultTwo four by four feature maps that differ only in where the single strong response sits. Max pooling over each quadrant returns the same value in both cases: the response survives, its position does not.max pooling, 2×21.01.0The strong response is in adifferent quadrant cell in eachinput, and the same after pooling.Bought: a small translation nolonger changes the answer.Paid: you cannot say where it was,which is why segmentation networkshave to put the resolution back.
    Fig. 79 — Two activations differing only in position, pooled to the same value. Invariance bought, location spent.

    FigCnnLineage

    Book II, Ch. II, Prop. 5permalink
    Five architectures, one questionFive convolutional architectures in order, with their depth in layers. Each is an answer to the same problem: how to add depth without the optimisation failing. ResNet's residual connection is the step that made depth cheap.how do you get deeper?5 layersLeNet1998it works at all8 layersAlexNet2012ReLU, dropout, GPUs19 layersVGG2014only 3×3, stacked152 layersResNet2015an additive path home66 layersEfficientNet2019scale the three axes togetherDepth axis is logarithmic. Before ResNet the fight was optimisation; after it, tuning.
    Fig. 80 — Five architectures asking one question: how to get deeper without the optimisation failing. The depth axis is logarithmic.

    FigFeatureHierarchy

    Book II, Ch. III, Prop. 1permalink
    What each depth responds toFour stages of a convolutional stack. Early layers respond to edges and colour, middle layers to texture and parts, late layers to whole objects. The early ones are generic across tasks, which is what makes transfer work.genericspecificlayer 1edges and colour blobslayer 2corners, texture, repeatslayer 3parts — an eye, a wheellayer 4objects and whole scenesNobody specified this. The hierarchy is what gradient descent produces with the loss at the far end.An edge is an edge whatever the task, which is why early layers transfer and late ones do not.
    Fig. 81 — What each depth responds to, from edges to whole objects. Nobody specified it; it is what the loss at the far end produces.

    FigTransferFreeze

    Book II, Ch. III, Prop. 2permalink
    How much of the network you let moveThree stacks of five blocks. In the first, four are frozen and only the head is trained. In the second, the upper half moves. In the third, everything moves. Each step down needs roughly an order of magnitude more labelled data than the one above it.frozentrainedlinear probehundreds of labelspartial fine-tunethousandsfull fine-tunetens of thousandsThe head isalways trained.Everything belowit is a decisionabout how muchdata you have.Fine-tuning a large backbone on four hundred images is how you get a model that memorises them.
    Fig. 82 — How much of the network you let move, and what each choice costs in labels.

    FigPretrainStart

    Book II, Ch. III, Prop. 3permalink
    Held-out error against labelled examplesTwo curves against the number of labelled examples. Training from scratch starts high and needs many labels. Starting from pretrained weights begins far lower and converges sooner. The architecture is identical in both.held-out error against labelled examplesfrom scratchfrom pretrained weightssame model,different startThe gap is widest on the left. Where labels are plentiful the two converge.
    Fig. 83 — Held-out error against labelled examples, from scratch and from pretrained weights. Same architecture, different starting point.

    FigPretextTask

    Book II, Ch. III, Prop. 4permalink
    Three pretext tasks built from one unlabelled imageAn unlabelled image feeds three tasks whose answers are known by construction: predict a hidden region, decide whether two crops came from the same photo, and recover an applied rotation. No annotator is involved.one unlabelled imageno labelhide a regionwhat was there?take two cropssame photo or not?rotate itby how much?The answer is known because you performed the corruption. What the model has to learn in order to undo it isthe representation you were after, and the choice of corruption decides which representation you get.
    Fig. 84 — Three tasks whose answers are known by construction, because you performed the corruption yourself.

    FigImageToPatches

    Book II, Ch. IV, Prop. 1permalink
    Image to patches to tokensAn image divided into a grid of fixed-size patches. Each patch is flattened into a vector and multiplied by a learned matrix to give an embedding. A positional embedding is added, and the result is an ordinary sequence of tokens for a transformer encoder.imagepatchesflatten + W+ positionencoderone vectorper patch+p1+p2+p3+p4+p5+p6No vision-specific machinery anywhere in it. Once an image is a sequence of tokens it concatenates with text.
    Fig. 85 — Image to patches to tokens. No vision-specific machinery anywhere in it, which is what made vision-language models straightforward.

    FigPatchSize

    Book II, Ch. IV, Prop. 2permalink
    Patch size against token count and attention costA 224 pixel image at three patch sizes. Halving the patch quadruples the number of tokens, and attention cost grows as the square of the token count, so it rises sixteenfold for each halving.224px imageViT-B/3249 tokensattention ×1ViT-B/16196 tokensattention ×16ViT-B/8784 tokensattention ×256Smaller patches mean finer detail and a quadratically larger bill. The /16 is that decision,and it is fixed at pretraining time — changing it later means re-interpolating the position embeddings.
    Fig. 86 — Patch size against token count. Halving the patch quadruples the tokens and multiplies attention cost by sixteen.

    FigPatchPosition

    Book II, Ch. IV, Prop. 3permalink
    The same patches, with and without positionA grid of patches and a shuffled copy of the same patches. To self-attention alone the two inputs are indistinguishable, because nothing in the operation references an index. The positional embedding added to each patch is what makes them different objects.to attention alone, these are the same inputas photographedshuffled+ p1+ p2+ p3+ p4+ p5position added to eachpatch before layer oneInterpolating these embeddings is how a model pretrained at 224px is evaluated at 448px.
    Fig. 87 — A grid of patches and a shuffled copy. To attention alone they are the same input until a positional embedding is added.

    FigSwinWindows

    Book II, Ch. IV, Prop. 4permalink
    Windowed attention, and the shift between blocksAn eight by eight grid of patches partitioned into four attention windows. In the next block the partition is offset by half a window, so patches that were sealed into separate windows now share one. Cost stays linear in the number of patches; information still crosses the whole image with depth.block k — windowsblock k+1 — shiftedsealed into four windowsthe same four patches now share oneAttention within a window costs the square of the window, not of the image — dense predictionaffordable. The shift is what stops the windows being four separate images.
    Fig. 88 — Windowed attention and the half-window shift between blocks. Without the shift the windows are four separate images.

    FigCollapse

    Book II, Ch. V, Prop. 1permalink
    Collapse, and what prevents itOn the left, embeddings spread across the space. In the middle, every embedding has fallen onto one point, which drives the agreement objective to zero while carrying no information. On the right, negatives push the embeddings apart and keep the space occupied.agreement onlythe degenerate optimumwith negativestrainloss is zeroand the model is useless"Make two views of thesame photo agree" issolved by a constant.Negatives supply therepulsion. So does astop-gradient, cheaper.
    Fig. 89 — The degenerate optimum: every input mapping to one point drives the agreement objective to zero and carries no information.

    FigAugmentationInvariance

    Book II, Ch. V, Prop. 2permalink
    Each augmentation declares an invarianceFour augmentations, each naming the factor the representation is being told to ignore. Colour jitter and greyscale discard hue, which is correct for natural images and destructive for a retinal photograph, where the colour of a lesion is the diagnosis.what you augment is what you discardrandom cropignore scale and framingflipignore left–rightcolour jitterignore illuminationgreyscaleignore hueon a fundus photographHaemorrhages are dark red.Exudates are yellow-white.Jitter the colour hard enoughand you have trained a modelto ignore the finding.The recipe that works on ImageNet encodes assumptions about photographs of objects. Every one of them isa claim about your domain, and they are rarely checked before being copied.
    Fig. 90 — Each augmentation names a factor the representation is told to ignore. On a fundus photograph, colour is the finding.

    FigStopGradient

    Book II, Ch. V, Prop. 3permalink
    Student and teacher, with the gradient cutTwo views of one image go through two branches. The student has a prediction head and is trained; the teacher is an exponential moving average of the student and receives no gradient. Cutting the gradient on one side is what stops both branches walking together into a constant.two views of one imageview Aview BstudenttrainedteacherEMA of the studentpredictor≈agreestop-gradientNo negatives anywhere. The teacher moves only through the moving average, so it is a slowly receding targetrather than a partner in the collapse. DINO adds centring and sharpening to the same skeleton.
    Fig. 91 — Student and teacher with the gradient cut on one side. The asymmetry is what replaces negatives.

    FigMaskedImage

    Book II, Ch. V, Prop. 4permalink
    Masked image modellingA grid of patches with a large fraction hidden. The encoder sees only the visible patches; a light decoder predicts what was removed. The training signal is the input itself, and the mask ratio is high because neighbouring patches are highly redundant.inputencoder seesdecoder predicts36 patches24 kept — 33%loss on the hidden ones onlyMask ratios of 75% work because neighbouring patches are redundant; hide too little and copying wins.
    Fig. 92 — Masked image modelling. High mask ratios work because neighbouring patches are redundant; hide too little and copying beats understanding.

    FigOneImageManyTasks

    Book II, Ch. VI, Prop. 1permalink
    One image, five tasksA single image and backbone feeding five different heads: classification, detection, segmentation, depth and retrieval. The backbone is the same in each case; what differs is the head and the shape of what comes out.one imageone backbonefive headsencoderunchangedclassificationone labeldetectionboxes + labelssegmentationa label per pixeldeptha distance per pixelretrievala vector to compareNothing about the pixels changed between these five rows. What changed is the question, and with itthe output and the loss that scores it — which is why a good backbone is worth more than a good head.Classification throws the spatial axes away. Everything below it has to keep them, or get them back.
    Fig. 93 — One image and one backbone feeding five heads. What changes is the question, the output shape and the loss.

    FigDetectionHeads

    Book II, Ch. VI, Prop. 2permalink
    Two questions, two lossesA detector must decide what an object is and where its box sits. The two are scored by different losses — a classification loss and a regression loss — and a prediction can be right about one and wrong about the other.one image, two predictionscat 0.94cat 0.41what is it?a classification lossover a fixed label setwhere is it?a regression loss on fourbox coordinates, usually IoUA box in the right place with the wrong label and the reverse fail differently,and mAP folds both into one number by sweeping the confidence threshold and the IoU threshold together.Reading a single mAP without knowing which threshold was used is reading a summary of a summary.
    Fig. 94 — Detection asks what and where, scored by two different losses. A prediction can be right about one and wrong about the other.

    FigSegDecoder

    Book II, Ch. VI, Prop. 3permalink
    The U, and why the skips are load-bearingAn encoder that halves the resolution at each stage and a decoder that doubles it back, joined by skip connections at every level. The bottleneck knows what is in the image; only the skips still know exactly where.resolution downand back upbottleneck — knows what, not whereskip connectionsPooling bought invariance by discarding position. A per-pixel task needs that position back, and it cannot berecovered from the bottleneck — so it is carried around it. Remove the skips and the boundaries go soft.
    Fig. 95 — The U and its skips. The bottleneck knows what is in the image; only the skips still know exactly where.

    FigDepthAmbiguity

    Book II, Ch. VI, Prop. 4permalink
    Two scenes, one projectionA camera centre with two objects along the same ray: a small one nearby and a large one far away. Both fill the same region of the image, so a single photograph cannot separate size from distance. Monocular depth is therefore predicted only up to an unknown scale.the same pixelscameranear, smallfar, largeimageBoth objects subtend the same angle, so both occupy the same pixels.Monocular depth models predict relative or scale-invariant depth — a metric claim needs stereo,a known object size, or a sensor that measures distance directly.
    Fig. 96 — Two scenes projecting to identical pixels. One photograph cannot separate size from distance.

    FigRetrievalSpace

    Book II, Ch. VI, Prop. 5permalink
    Retrieval is a metric, not a label setAn embedding space with three clusters. A query lands in it and the answer is its nearest neighbours. Nothing about the procedure requires the classes to have existed when the model was trained.embedding spacequeryThe output is an ordering,not a class. Add a newcategory tomorrow andnothing needs retraining —you index one more vector.Which is why the evaluation is recall@k and mAP rather than accuracy, and why the embedding's geometry —not its accuracy on any fixed label set — is the thing that has to be good.
    Fig. 97 — Retrieval returns an ordering, not a class, so a category that did not exist at training time costs one more indexed vector.

    FigModalities

    Book II, Ch. VII, Prop. 1permalink
    Four modalities and what each one measuresA table of four imaging modalities with the physical quantity each measures and the units it reports. CT has absolute units and can be windowed by value; MRI has none, so its intensities are only comparable within a scan.modalitymeasuresunitsconsequenceX-rayattenuation, projectedarbitraryeverything overlapsCTattenuation, reconstructedHounsfield unitsabsolute and comparableMRIproton relaxationnonescanner-dependent intensitiesFundusreflected lightRGBcolour is the findingMin-max normalising a CT throws away the property that makes it a measurement. Not normalising anthe scanner in the signal. The right preprocessing is a fact about the physics, not a default in a library.
    Fig. 98 — Four modalities and what each measures. The right preprocessing is a fact about the physics, not a default in a library.

    FigRocSweep

    Book II, Ch. VII, Prop. 2permalink
    The ROC curve is a sweep, not a scoreA receiver operating characteristic curve, traced from the strict end to the permissive end. Each point on it is one decision threshold with its own sensitivity and specificity. AUROC summarises the whole curve; a deployed system occupies a single point on it.sensitivity against 1 − specificitychancestrict — few alarmspermissive — few missesAUROCis thearea underall of itScreening wants the permissive end and pays in false alarms. Confirmation wants the strict end.
    Fig. 99 — The ROC curve traced from strict to permissive. Every point is a threshold; a deployed system occupies exactly one of them.

    FigOperatingPoint

    Book II, Ch. VII, Prop. 3permalink
    The same model at two prevalencesOne model with 90% sensitivity and 90% specificity applied to a thousand patients. At ten percent prevalence, half of its positive calls are correct. At one percent, fewer than one in ten are. Sensitivity and specificity did not change.sensitivity 90% · specificity 90% · 1000 patientsprevalence 10%of the positive calls…90 true90 falseprecision 50%prevalence 1%of the positive calls…9 true99 falseprecision 8%Sensitivity and specificity are properties of the model. Precision is a property of the model andso a screening tool validated on an enriched cohort will disappoint in a clinic where the disease is rare.Ask what the prevalence was in the test set before reading any precision or F1.
    Fig. 100 — The same sensitivity and specificity at two prevalences. Precision is a property of the model and the population together.

    FigCalibrationCurve

    Book II, Ch. VII, Prop. 4permalink
    Predicted probability against observed frequencyA reliability diagram. The diagonal is perfect calibration. The plotted curve sits below it, meaning that among cases the model calls ninety percent likely, far fewer than ninety percent are positive. Its ranking may still be excellent.observed frequency against predicted probabilityperfectly calibratedoverconfidentsays 0.9is 0.62AUROC is unchanged by any monotone rescaling of the scores, so it cannot see this at all.
    Fig. 101 — Predicted probability against observed frequency. AUROC is unchanged by any monotone rescaling, so it cannot see this at all.

    FigDomainShift

    Book II, Ch. VII, Prop. 5permalink
    The same model on three populationsOne model evaluated on an internal held-out set, on data from a different scanner at the same hospital, and on data from another institution. The internal number is the one that appears in the abstract and the external one is the one that predicts deployment.the same weights, three populationsinternal test0.94same hospital, new scanner0.87external hospital0.71Everything upstream can be right — the split, the calibration, the augmentation — and this gap opens,because every one of those checks drew from the same distribution the model was fitted on.The only measurement that answers it is data from an institution outside the training set.
    Fig. 102 — One model on three populations. Every internal check drew from the distribution the model was fitted on.

    FigVisionMap

    Book II — closingpermalink
    The whole of Book II on one plateA vertical map from raw pixels to a decision. Pixels are normalised, mixed either by convolution or by attention over patches, built into a hierarchy of features, read by a task head, and finally turned into a number that someone acts on. Each stage is labelled with the chapter that covers it.from pixels to a decisionpixels, channels, normalisationCh. Iconvolution or patchesCh. II, IVa learned hierarchy of featuresCh. III, Va head for the question askedCh. VIa number somebody acts onCh. VIIone imageTwo stages are architecture and three are judgement. Most of the literature is about the second row;goes wrong in a deployed system is in the last one.If a stage here is not one you could explain to somebody else, the chapter is named beside it.
    Fig. 103 — The whole of Book II on one plate. Two stages are architecture and three are judgement.