Elementa
The figure library Every diagram on this site, at its own address. All are hand-written inline SVG using the site palette, so they stay legible in both themes and paste into a slide deck without becoming a screenshot of a screenshot.
microaneurysms haemorrhages hard exudates cotton wool spots vessels
FigPipelineHierarchiRetina
The HierarchiRetina three-stage pipeline A fundus image enters Stage I screening for referable diabetic retinopathy at grade 2 or above. It then branches to five parallel lesion segmentation models — microaneurysms, haemorrhages, hard exudates, cotton wool spots and vessels — whose masks rejoin the raw image as eight channels into Stage III, the lesion-guided grader. Stage III emits either an ordinal severity grade from one to four, or an ungradable verdict routed to human review. Fundus image Stage I · screening (referable DR, grade ≥ 2) MA HSMoE HE MultiScale EX BrightSpot CWS warm-up Vessels FOV mask Stage III · LG-DRG (8 channels: RGB + 5 masks) Severity grade (CORN ordinal, 1–4) Ungradable (routed to human review)
Fig. 1 — The HierarchiRetina three-stage pipeline: screening, five parallel lesion segmenters, and lesion-guided grading over eight channels.Fig. 2 — A model as a function of two arguments. The input comes from the world; the parameters are the only part training is allowed to change.Fig. 3 — The axes of a tensor, named. The same numbers arranged along different axes are different data, which is why a shape is a claim about meaning and not only about storage.Fig. 4 — Training as a path through parameter space. Every point on the path is the same architecture holding different numbers.Fig. 5 — A loss surface drawn as contours, with a deep basin and a shallower one. Descent finds the nearer minimum, which is not always the better one.Fig. 6 — Training error and held-out error against training time. The gap that opens after the held-out minimum is the memorised part.Fig. 7 — The weight vector is the normal to the boundary, so the bias slides the boundary along w and can never turn it.The score divided by the norm of w is a signed distance A point sits away from the boundary. The perpendicular dropped from it to the boundary has length equal to the score divided by the length of the weight vector; the sign of the score says which side, and the magnitude says how far. one number, two facts boundary x |⟨w,x⟩+b| / ‖w‖ x′ sign → the class size → the distance scores rescale with w; distances do not doubling w and b doubles every score and moves nothing
Fig. 8 — One number carries both facts: its sign is the predicted class and its magnitude, once divided by the norm of w, is the distance to the boundary.Fig. 9 — An update turns the weight vector toward the point it got wrong, and the boundary swings with it. A point already correct produces no change at all.XOR puts each class on one diagonal of the square, and the diagonals cross The four corners of the unit square carry XOR's labels. The two positive corners lie on one diagonal and the two negative corners on the other. Any straight line separating the ends of one diagonal must pass between the ends of the other, so no line can separate the classes. XOR on the unit square 0,0 → 0 1,1 → 0 1,0 → 1 0,1 → 1 b ≤ 0 w₂ + b > 0 w₁ + b > 0 w₁ + w₂ + b ≤ 0 add the middle two: b > 0 against the first: impossible filled = 1 · hollow = 0 no straight line separates the filled corners from the hollow ones
Fig. 10 — XOR puts each class on one diagonal of the square, and the diagonals cross. Any line separating one diagonal must pass between the other.Fig. 11 — The left panel is the number you report; the right panel is what moves the model. Huber tracks the parabola in value and caps its pull at δ.Fig. 12 — Choosing the loss on the right asserts the distribution on the left, whether or not the assertion was intended.Cross-entropy cancels the output derivative; squared loss keeps it Two chains from a logit to a gradient. Along the top, cross-entropy contributes a factor that is the reciprocal of the sigmoid derivative, so the two cancel and the gradient is simply p minus y. Along the bottom, squared loss contributes no such factor, so the sigmoid derivative survives and shrinks the gradient toward zero exactly where the model is most wrong. cross-entropy squared loss ∂L/∂p (p−y) / p(1−p) × dp/dz p(1−p) = p − y these cancel exactly bounded by 1 largest when most wrong ∂L/∂p p − y × dp/dz p(1−p) = (p−y)·p(1−p) nothing cancels → 0 as the model worsens at z = −10 with y = 1: cross-entropy sends −0.99996, squared loss sends −0.0000454
Fig. 13 — Cross-entropy contributes the reciprocal of the output derivative, so the two cancel. Squared loss contributes nothing, so the derivative survives and shrinks the gradient exactly where the model is most wrong.Fig. 14 — The sigmoid derivative never exceeds one quarter, and a product of them reaches the smallest half-precision number by depth twelve — in the best case that never occurs.Fig. 15 — A dead ReLU unit is a fixed point of gradient descent: the derivative that would move its weights is the same zero that killed it.Fig. 16 — ReLU and GELU agree away from the origin and differ only in a narrow band around it — which is exactly where the gradient is decided.Fig. 17 — One unit: a weighted sum, a bias, and an activation. Everything before the activation is affine, and affine maps compose into a single affine map.Fig. 18 — Two linear layers composed give one line whatever the depth. The same two layers with a rectifier between them give a piecewise curve, one fold per unit.Fig. 19 — One graph traversed twice: values forwards, partial derivatives backwards. Backpropagation is the chain rule with the intermediate results kept.Fig. 20 — The same descent under three learning rates. Too small and it creeps, too large and it climbs the far wall; the gradient supplies the direction, never the distance.Fig. 21 — Seven points fitted exactly and fitted loosely. Regularisation does not improve the fit — it decides which fit you get when many are available.Batch and layer normalisation reduce over different axes Two identical three-by-three matrices have examples in rows and features in columns. Batch normalisation highlights a column: one feature across three examples. Layer normalisation highlights a row: three features within one example. The other columns or rows form their own separate groups. Neither operation changes the tensor shape. batch normalisation features → 1 4 7 2 5 8 3 6 9 examples one feature, across examples Other examples affect this output. Inference normally freezes the statistics. layer normalisation features → 1 4 7 2 5 8 3 6 9 examples one example, across features Other examples do not affect it. The same rule at training and inference.
Fig. 22 — The same matrix, different reduction groups. Batch normalisation shares statistics down a feature column; layer normalisation shares them across one example. Highlighted entries form one group; the remaining rows or columns form their own groups.One string under three tokenisations The string "unhappiness" segmented three ways: as subwords un / happi / ness, as a single word token, and as ten individual characters. Each segmentation fixes which distinctions the model is able to represent at its input. unhappiness subword un happi ness word unhappiness char un h a p p i n e s s
Fig. 23 — One string under three tokenisations. The segmentation chosen at the input fixes which distinctions the model is able to represent at all.Byte-pair encoding, three merges deep The word lowest, first split into single characters, then merged three times. Each row shows the vocabulary after one merge of the most frequent adjacent pair, ending with the two pieces low and est. most frequent adjacent pair, merged start l o w e s t merge 1 l o w e st merge 2 lo w e st merge 3 low est vocabulary: 6 + "st" → 7 + "lo" → 8 + "low", "est" → 10 The merge list is the tokeniser. Nothing in it is learned, and it is fixed before training begins.
Fig. 24 — Byte-pair encoding, three merges deep. The merge list is the tokeniser; it is fixed by frequency in a corpus before training begins, and nothing in it is learned by the model.Fig. 25 — The embedding table and the space it induces. Retrieval is indexing; the geometry among the retrieved rows is a consequence of training.Fig. 26 — One token in two sentences. It enters as the same vector in both, and only separates once a layer of attention has let it read its neighbours.Fig. 27 — A recurrence unrolled. Every step passes a fixed-width state to the next, so the path between two distant positions is as long as the distance between them.Fig. 28 — A repeated multiplication over sixteen steps. Nothing here is peculiar to recurrence: it is what happens to any long product of numbers that are not exactly one.Fig. 29 — A gated cell. The carry line crosses only a multiplication and an addition, so with the forget gate near one the gradient has a route home that no weight matrix attenuates.Receptive field growth under stacked dilated convolutions Four rows of positions. Each layer reads three positions from the row below at a spacing that doubles, so the window seen by the top unit widens from three positions to fifteen across three layers. input layer 1 dil 1 · sees 3 layer 2 dil 2 · sees 7 layer 3 dil 4 · sees 15 Range grows exponentially with depth, and every position is computed at once.
Fig. 30 — Receptive field growth under stacked dilated convolutions. Range grows exponentially with depth and every position is computed at once — attention buys the same range at depth one.Four ways to read a sequence, by path length and parallelism A table of four architectures. Recurrent models connect distant positions through a path that grows with distance and cannot be parallelised. Convolution shortens the path logarithmically and parallelises. Attention connects any two positions in one step, in parallel, at quadratic cost. architecture path between two positions over the sequence RNN O(n) sequential LSTM / GRU O(n) sequential CNN O(log n) parallel Attention O(1) parallel Attention is not a better idea than recurrence in the abstract. It is the trade that pays when the hardware is parallel and the sequences are short enough to afford n².
Fig. 31 — Four ways to read a sequence, placed by path length and parallelism. Attention is not a better idea in the abstract; it is the trade that pays on parallel hardware.Attention as a weighted average over values A single query is compared against four keys. The resulting softmax weights — 0.06, 0.61, 0.09 and 0.24 — are shown as horizontal bars, and the output is the sum of the value vectors scaled by those weights. The weights come from content, not from position. query q softmax(q · kᵢ / √d) k₁ 0.06 k₂ 0.61 k₃ 0.09 k₄ 0.24 Σ wᵢ vᵢ weights are content-addressed — nothing here depends on i
Fig. 32 — Attention as a weighted average over values, with weights computed by comparing a query against every key. Nothing in the computation depends on position.Fig. 33 — Query, key and value as three learned projections of one vector: what a token is looking for, what it advertises, and what it hands over once matched.The same attention logits, scaled and unscaled Two bar charts of attention weights over eight keys. Scaled, the weights are spread across several keys. Unscaled, almost all the weight lands on one key and the remaining bars are invisible, which is where the gradient dies. scaled by 1/√d k1 k2 k3 k4 k5 k6 k7 k8 max weight 27% unscaled — logits grow with d k1 k2 k3 k4 k5 k6 k7 k8 max weight 80% For unit-variance q and k the dot product has variance d, so the logits grow with the head width. Dividing by √dₖ restores it, whatever the width.
Fig. 34 — The same attention logits, scaled and unscaled. Without the divisor the distribution collapses onto one key and the gradient through the softmax goes flat.Fig. 35 — Self-attention and cross-attention are one operation under two wirings. Only the source of the keys and values changes.The attention score matrix at three sequence lengths Three square grids of scores, for sequences of four, eight and sixteen tokens. Each doubling of the sequence length quadruples the number of scores, because every token is compared against every other. one score per pair of positions n = 4 16 scores n = 8 64 scores n = 16 256 scores Path length is constant — any token reaches any other in one step — and the square is what it costs.
Fig. 36 — The score matrix at three sequence lengths. Constant path length between any two positions is bought with an area that grows as the square of the sequence.Fig. 37 — Permutation equivariance and its repair. Without positional encoding a reordered input yields a merely reordered output; with it, order becomes information.Fig. 38 — Sinusoidal encoding as a bank of clocks. Fast dimensions separate neighbouring positions, slow ones separate distant regions, and together they give each position a signature.Rotary position embedding A circle with three vectors drawn at angles proportional to their positions in the sequence. Because both the query and the key are rotated, the angle between any two of them depends only on how far apart the positions are. position becomes an angle m = 2 m = 5 m = 9 ⟨ R(mθ)q , R(nθ)k ⟩ = ⟨ q , R((n−m)θ)k ⟩ The absolute indices cancel. What survives the inner product is n − m, the distance. Nothing is added to the embedding, so the residual stream is left alone; only q and k are turned, inside each head.
Fig. 39 — Rotary encoding turns position into an angle. Both the query and the key are rotated, so the absolute indices cancel in the inner product and only the distance survives.Fig. 42 — One attention budget divided into four heads. Heads are slices rather than copies, so several relations are attended to at once at the cost of resolution inside each.Fig. 43 — The residual stream. Blocks read from it and add back into it rather than replacing it, which is both why gradients reach layer one and why intermediate states can be read.Fig. 44 — Post-norm puts a normalisation on the shortcut; pre-norm leaves the shortcut clean. The same two components in two orders, with different training behaviour.The causal mask A square grid of attention scores for seven tokens. The lower triangle, including the diagonal, is open: each token may attend to itself and to everything before it. The upper triangle is closed, because those positions lie in the future. query row attends to key column the −∞ −∞ −∞ −∞ −∞ −∞ cat −∞ −∞ −∞ −∞ −∞ sat −∞ −∞ −∞ −∞ on −∞ −∞ −∞ the −∞ −∞ warm −∞ mat the cat sat on the warm mat Set before the softmax, so the masked entries receive exactly zero weight rather than a small one. This is what lets every position be trained at once on the same sequence: n prediction problems, one forward pass. Remove the triangle and the same weights become an encoder. Nothing else about the block changes.
Fig. 45 — The causal mask. Setting the future to minus infinity before the softmax is what lets every position in a sequence be trained as a separate prediction in one pass.Fig. 46 — Encoder, decoder, and the two joined by cross-attention. Same block, same arithmetic — the masks and the wiring are the whole of the difference.Fig. 47 — The batch gradient as an estimate of the true one. A small batch scatters widely and steps often; a large batch points true and steps rarely.Fig. 48 — Warm-up then decay. The rise exists because the optimiser has no statistics yet; the fall exists because a large step near the end discards what the earlier ones found.Weight decay and gradient clipping guard different things A bar chart of gradient norms over fourteen steps, mostly small with two enormous spikes. A horizontal clipping threshold cuts only the two spikes. Beneath, a separate note shows weight decay acting on every step regardless. gradient norm per step clip threshold Clipping is a rare intervention: on almost every step it does nothing at all. Weight decay is the opposite: a small constant pull toward zero, applied on every step to every parameter, whatever the gradient is doing. One bounds the step, the other bounds where they land.
Fig. 49 — Gradient norms over fourteen steps with a clipping threshold. Clipping touches only the rare spike; weight decay pulls on every parameter every step. They guard different failures.Representable range in three float formats Three horizontal bands on a logarithmic magnitude axis. Half precision covers a narrow band; brain float and single precision cover a wide one. A marked region of small gradient magnitudes falls below the half-precision floor and rounds to zero. magnitude, log scale fp32 fp16 bf16 where gradients live late in training Loss scaling multiplies the loss by a large constant so the gradients land inside the band, then divides it back out before the update.
Fig. 50 — Representable range in three float formats. Half precision is narrow, and late-training gradients fall through its floor — which is what loss scaling exists to prevent.Two ways to split the same six images Six images from three patients, two each. Split by image, every patient appears on both sides of the split and the held-out score is inflated. Split by patient, no patient appears on both sides and the score is honest. split by image train p1 p2 p3 held out p1 p2 p3 patients 1, 2 and 3 all appear on both sides measures recall of patients already seen split by patient train p1 p1 p2 p2 held out p3 p3 no patient appears on both sides measures what happens on a new patient The unit of the split has to be the unit the claim is about — patient, site, scanner, study. Get it wrong and every other measure of rigour is decoration: the number was decided before training began.
Fig. 51 — The same six images split two ways. Splitting by image puts every patient on both sides; splitting by patient is the only one of the two that measures what the claim is about.FigGramAnchoring
Marginalia II — DINOv2 and DINOv3 permalink copy SVG Fig. 53 — Dense features rot as a large ViT trains, because the image-level objective wants abstraction and wins. Anchoring the patch-to-patch Gram matrix to an earlier checkpoint drags them back.FigFcmaeGrn
Marginalia III — ConvNeXt V2 permalink copy SVG Fig. 54 — A kernel slides across a masked hole and leaks the answer; sparse convolution stops it, and global response normalisation stops the channels collapsing. Neither half works alone.FigConvHandback
Marginalia IV — Swin UNETR V2 permalink copy SVG Fig. 55 — One residual convolution at the head of every encoder stage. That single block is the entire contribution of the paper.FigLinearProbe
Marginalia V — logistic regression permalink copy SVG Fig. 56 — What a linear probe measures, and what does the measuring. The encoder is frozen; the classifier on top is logistic regression, and its convexity is why it can serve as a ruler at all.Dataset fingerprint to pipeline Five measured properties of a dataset enter a set of heuristic rules, which resolve them jointly under a GPU memory budget into five pipeline decisions. The architecture is not among the things being chosen. fingerprint pipeline voxel spacings image sizes intensity distribution modality class ratios heuristic rules under a memory budget target spacing resampling normalisation patch and batch size depth and pooling Fixed regardless: Dice plus cross-entropy, SGD at 0.99, 1000 epochs, the same augmentations every time. The original paper used no residual connections, no attention, no squeeze-and-excitation. That was the claim.
Fig. 57 — A dataset fingerprint resolved by heuristic rules into a whole pipeline under a memory budget. The architecture is not among the things being chosen.Zero-shot average across 27 benchmarks Three bars on a truncated axis. Going from eight billion parameters to eighteen billion gains seven tenths of a point. Raising the input resolution on the smaller model, with no extra parameters at all, gains six. zero-shot average, 27 benchmarks 79.4% EVA-CLIP-8B 224px 80% EVA-CLIP-8B 448px, same params 80.7% EVA-CLIP-18B 224px, 2.2x params 78.5 Axis truncated. The data stayed fixed the whole way — scaling parameters still works, it just works slowly, and the lever everyone else is pulling is the data.
Fig. 58 — Doubling the parameters on a fixed dataset bought seven tenths of a point. Raising the input resolution, with no extra parameters at all, bought six.The open ladder and the closed rung Five rungs of a model ladder, each wider than the last. The four lower ones ship as open weights. The top rung stayed API-only for most of the family's life, and the gap between the two is the measure of the commitment. ship the whole ladder at once laptop open weights workstation open weights server open weights cluster open weights Max class API only, until Aug 2026 Watch what they hold back, not what they release. That gap has moved in both directions.
Fig. 59 — The whole ladder shipped at once, with the top rung held back. The gap between the open rungs and the closed flagship is the honest measure of the commitment.The two checkpoints in one release A comparison of the 2.4 trillion parameter flagship against the 27 billion parameter model released beside it. The headline belongs to the first; the second is the one most people can actually run. Qwen3.8-2.4T-A95B Qwen3.8-27B parameters 2.4T, 95B active 27.8B dense input text only multimodal thinking always on switchable licence bespoke, revenue clauses Apache 2.0 you can host it no yes The open checkpoint is not the API model, and the licence is not Apache whatever the headlines said.
Fig. 60 — The two checkpoints in one release. The headline belongs to the 2.4 trillion; the 27B is the one that fits on hardware you can buy, and it is the Apache one.Fig. 61 — Whether the concept is present and where it is are different questions that used to interfere. Giving presence its own token and multiplying the scores is most of why the numbers moved.Cost per token at a one-million-token context Two bars per model. Compared with the previous generation, V4 needs about twenty-seven percent of the per-token inference operations and ten percent of the key-value cache at a one-million-token context. at 1M context, relative to V3.2 V3.2 FLOPs 100% KV cache 100% V4 FLOPs 27% KV cache 10% Compressed Sparse Attention interleaved with Heavily Compressed Attention. Everyone quoted the 1.6 trillion. The number that matters is the 10%.
Fig. 62 — At a one-million-token context, roughly a quarter of the per-token operations and a tenth of the key-value cache. Everyone quoted the parameter count.Rate card against cost per task Three models compared twice: by their published price per million input tokens, and by measured cost to complete a task. The ordering is not the same, because a model whose answers run token-efficient pays less per job than its rate card suggests. price per million in measured cost per task Kimi K3 $3.00 $0.94 GPT-5.6 Sol $2.20 $1.04 Opus 4.8 $4.50 $1.80 Three times the rate of K2.6 and a lower bill than either rival. The question is not whether K3 is better. It is whether your workload notices.
Fig. 63 — Rate card against measured cost per task. A higher price per million tokens and a lower bill per job are not contradictory when the answers run token-efficient.FigTotalVsActive
Marginalia XVII — Mixture of experts permalink copy SVG Fig. 64 — Every expert is stored; only the routed few compute. Memory tracks the big number and arithmetic tracks the small one, which is the whole reason anyone does this.Long-context comprehension at 120K tokens Accuracy on long-context comprehension at a hundred and twenty thousand tokens. Scout, which advertises a ten million token window, scores fifteen point six percent, well under a competitor at the same length. comprehension at 120K tokens Gemini 2.5 Pro 90.6% Maverick 28.1% Scout 15.6% Scout advertises 10,000,000 tokens. Needle-in-a-haystack retrieval across that window is genuinely strong, and retrieval is not comprehension.
Fig. 65 — Comprehension at 120K tokens against an advertised ten-million-token window. Retrieval across a context is not comprehension of it.FigFlopsVsLatency
Marginalia XX — EfficientNet permalink copy SVG Fewer operations, more time, same accuracy Two models that both reach 84.0% on ImageNet. One uses 1.8 times fewer floating-point operations and runs 2.7 times slower on the same hardware, because depthwise convolutions are bound by memory bandwidth rather than arithmetic. both land at 84.0% on ImageNet EfficientNet-B6 FLOPs arithmetic time wall clock ResNet-RS-350 FLOPs arithmetic time wall clock Accelerators are bound by memory bandwidth, not arithmetic. "Efficient" meant FLOP-efficient, and a decade of practitioners read it as fast.
Fig. 66 — Two models at the same accuracy: one with 1.8x fewer operations and 2.7x more wall-clock time. Accelerators are bound by memory bandwidth, not arithmetic.FigContextResolution
Marginalia XXII — AlphaGenome permalink copy SVG Context length against output resolution Two axes: sequence context and output resolution. Earlier models occupy one corner or the other — fine resolution over a short window, or a long window with binned outputs. AlphaGenome occupies the corner that was assumed unreachable. context resolution ↑ · context → SpliceAI, BPNet 10kb, base resolution Enformer, Borzoi 200–500kb, 32–128bp bins AlphaGenome 1Mb, base resolution the corner nobody had The constraint was computational rather than conceptual. Eleven output modalities at once, from one sequence.
Fig. 67 — Context against output resolution. Every earlier model sits on one edge or the other; the claim is that the corner between them was an engineering limit rather than a law.Structure, affinity, and the speed between them Co-folding answers whether two molecules fit and runs fast. Free energy perturbation answers how tightly they hold and runs slowly enough to be a scheduling decision. Boltz-2 reaches comparable correlation to the second at the speed of the first. what each method answers co-folding can they fit fast FEP how tightly slow Boltz-2 how tightly fast Pearson 0.62, >1000x faster Approaching FEP is not matching FEP. A correlation of 0.62 supports ranking a library and choosing what to synthesise. It does not support telling a chemist that compound 47 binds at 12 nanomolar. Knowing which decisions a number can carry is most of the skill in using it.
Fig. 68 — Co-folding says whether two molecules fit, free energy perturbation says how tightly and runs slowly. Reaching the second at the speed of the first is what changes screening.Fig. 69 — Accuracy against pretraining set size. The paper's finding is a crossover with a threshold in it, somewhere near a hundred million images — not a verdict.Where scVI departs from a textbook autoencoder Three rows comparing an ordinary variational autoencoder with scVI: the likelihood is a negative binomial rather than a Gaussian, batch identity conditions both encoder and decoder rather than being corrected afterwards, and library size gets its own latent instead of contaminating the cell state. textbook VAE scVI count likelihood Gaussian, squared error negative binomial batch corrected afterwards conditioned on, in both halves sequencing depth left in the latent its own scaling factor Any method that ignores these produces a beautiful embedding of your experimental logistics. Three independent evaluations, different teams, different data, put this eight-year-old VAE ahead.
Fig. 70 — Three departures from a textbook autoencoder, each one a fact about the assay: a count likelihood, batch as a conditioning variable, and library size given its own latent.Two evaluations, opposite conclusions Arc reports that STATE is the first model in its domain to consistently beat simple linear baselines. VCBench reports that pre-registered baselines match or exceed every foundation model tested on four of five dimensions. Four ordinary differences account for the gap. Arc, June 2025 first to consistently beat simple linear baselines VCBench, June 2026 baselines match or exceed all five models on four of five both in print different baselines different splits different metrics a year apart Neither claim is dishonest. Working out how they can both be true is more useful than picking a side. If a benchmark cannot tell interaction structure from main effects, a good score on it demonstrates little.
Fig. 71 — Two credible evaluations pointing opposite ways, and the four ordinary differences that account for it. Neither claim is dishonest.Operators by the range they cover Four operator types striped through the model, each covering a different range: short and medium convolutions for local motifs, long implicit convolutions for domain-scale structure, and self-attention reserved for the sparse long-range relationships that need it. one megabase, single-nucleotide resolution short explicit conv motifs, splice sites medium regularized conv local regulatory syntax long implicit conv domain-scale structure self-attention enhancer to promoter, sparse Attention costs the square of the length, so you do not pay it across the whole sequence.
Fig. 72 — Operators assigned by the range they cover, so attention is spent only on the sparse long-range relationships that need it. A megabase becomes affordable.Fig. 73 — An image as a stack of channel planes. The numbers are only numbers; the shape is the claim about what they are.Fig. 74 — The same pixel values in a different arrangement. Every summary statistic is identical and one of the two is a picture.Intensity distribution before and after normalisation A histogram of pixel intensities shifted off centre and narrow, and the same histogram after subtracting the mean and dividing by the standard deviation. The statistics used must come from the training set, and the same ones must be applied at inference. raw intensities after (x − μ) / σ 0 0 μ and σ come from the training set, and the same two numbers must be applied at inference. Recomputing them per batch at test time leaks information between test examples, and the leak is silent. A model trained on one scanner's intensity distribution has been told what to expect. Another scanner is a different promise.
Fig. 75 — Intensities before and after normalisation. The statistics must come from the training set and be applied unchanged at inference.Fig. 76 — One kernel applied at every position. Nine weights and a bias, whatever the size of the image — the reuse is the locality prior.Output size under three settings A seven by seven input under three configurations of a three by three kernel. Without padding the output shrinks to five. With padding of one it keeps its size. With a stride of two it halves. 7×7 input, 3×3 kernel no padding, stride 1 output 5 × 5 padding 1, stride 1 output 7 × 7 padding 1, stride 2 output 4 × 4 out = ⌊(in + 2·pad − k) / stride⌋ + 1. Nothing else decides it, and it is the commonest shape bug.
Fig. 77 — A 7×7 input under three settings. Padding and stride decide the output size and nothing else does.Fig. 78 — What one unit at the top can see, opening by two positions a layer. Range is bought with depth.Fig. 79 — Two activations differing only in position, pooled to the same value. Invariance bought, location spent.Five architectures, one question Five convolutional architectures in order, with their depth in layers. Each is an answer to the same problem: how to add depth without the optimisation failing. ResNet's residual connection is the step that made depth cheap. how do you get deeper? 5 layers LeNet 1998 it works at all 8 layers AlexNet 2012 ReLU, dropout, GPUs 19 layers VGG 2014 only 3×3, stacked 152 layers ResNet 2015 an additive path home 66 layers EfficientNet 2019 scale the three axes together Depth axis is logarithmic. Before ResNet the fight was optimisation; after it, tuning.
Fig. 80 — Five architectures asking one question: how to get deeper without the optimisation failing. The depth axis is logarithmic.Fig. 81 — What each depth responds to, from edges to whole objects. Nobody specified it; it is what the loss at the far end produces.How much of the network you let move Three stacks of five blocks. In the first, four are frozen and only the head is trained. In the second, the upper half moves. In the third, everything moves. Each step down needs roughly an order of magnitude more labelled data than the one above it. frozen trained linear probe hundreds of labels partial fine-tune thousands full fine-tune tens of thousands The head is always trained. Everything below it is a decision about how much data you have. Fine-tuning a large backbone on four hundred images is how you get a model that memorises them.
Fig. 82 — How much of the network you let move, and what each choice costs in labels.Fig. 83 — Held-out error against labelled examples, from scratch and from pretrained weights. Same architecture, different starting point.Three pretext tasks built from one unlabelled image An unlabelled image feeds three tasks whose answers are known by construction: predict a hidden region, decide whether two crops came from the same photo, and recover an applied rotation. No annotator is involved. one unlabelled image no label hide a region what was there? take two crops same photo or not? rotate it by how much? The answer is known because you performed the corruption. What the model has to learn in order to undo it is the representation you were after, and the choice of corruption decides which representation you get.
Fig. 84 — Three tasks whose answers are known by construction, because you performed the corruption yourself.Fig. 85 — Image to patches to tokens. No vision-specific machinery anywhere in it, which is what made vision-language models straightforward.Patch size against token count and attention cost A 224 pixel image at three patch sizes. Halving the patch quadruples the number of tokens, and attention cost grows as the square of the token count, so it rises sixteenfold for each halving. 224px image ViT-B/32 49 tokens attention ×1 ViT-B/16 196 tokens attention ×16 ViT-B/8 784 tokens attention ×256 Smaller patches mean finer detail and a quadratically larger bill. The /16 is that decision, and it is fixed at pretraining time — changing it later means re-interpolating the position embeddings.
Fig. 86 — Patch size against token count. Halving the patch quadruples the tokens and multiplies attention cost by sixteen.Fig. 87 — A grid of patches and a shuffled copy. To attention alone they are the same input until a positional embedding is added.Windowed attention, and the shift between blocks An eight by eight grid of patches partitioned into four attention windows. In the next block the partition is offset by half a window, so patches that were sealed into separate windows now share one. Cost stays linear in the number of patches; information still crosses the whole image with depth. block k — windows block k+1 — shifted sealed into four windows the same four patches now share one Attention within a window costs the square of the window, not of the image — dense prediction affordable. The shift is what stops the windows being four separate images.
Fig. 88 — Windowed attention and the half-window shift between blocks. Without the shift the windows are four separate images.Fig. 89 — The degenerate optimum: every input mapping to one point drives the agreement objective to zero and carries no information.FigAugmentationInvariance
Fig. 90 — Each augmentation names a factor the representation is told to ignore. On a fundus photograph, colour is the finding.Fig. 91 — Student and teacher with the gradient cut on one side. The asymmetry is what replaces negatives.Masked image modelling A grid of patches with a large fraction hidden. The encoder sees only the visible patches; a light decoder predicts what was removed. The training signal is the input itself, and the mask ratio is high because neighbouring patches are highly redundant. input encoder sees decoder predicts 36 patches 24 kept — 33% loss on the hidden ones only Mask ratios of 75% work because neighbouring patches are redundant; hide too little and copying wins.
Fig. 92 — Masked image modelling. High mask ratios work because neighbouring patches are redundant; hide too little and copying beats understanding.Fig. 93 — One image and one backbone feeding five heads. What changes is the question, the output shape and the loss.Two questions, two losses A detector must decide what an object is and where its box sits. The two are scored by different losses — a classification loss and a regression loss — and a prediction can be right about one and wrong about the other. one image, two predictions cat 0.94 cat 0.41 what is it? a classification loss over a fixed label set where is it? a regression loss on four box coordinates, usually IoU A box in the right place with the wrong label and the reverse fail differently, and mAP folds both into one number by sweeping the confidence threshold and the IoU threshold together. Reading a single mAP without knowing which threshold was used is reading a summary of a summary.
Fig. 94 — Detection asks what and where, scored by two different losses. A prediction can be right about one and wrong about the other.Fig. 95 — The U and its skips. The bottleneck knows what is in the image; only the skips still know exactly where.Two scenes, one projection A camera centre with two objects along the same ray: a small one nearby and a large one far away. Both fill the same region of the image, so a single photograph cannot separate size from distance. Monocular depth is therefore predicted only up to an unknown scale. the same pixels camera near, small far, large image Both objects subtend the same angle, so both occupy the same pixels. Monocular depth models predict relative or scale-invariant depth — a metric claim needs stereo, a known object size, or a sensor that measures distance directly.
Fig. 96 — Two scenes projecting to identical pixels. One photograph cannot separate size from distance.Retrieval is a metric, not a label set An embedding space with three clusters. A query lands in it and the answer is its nearest neighbours. Nothing about the procedure requires the classes to have existed when the model was trained. embedding space query The output is an ordering, not a class. Add a new category tomorrow and nothing needs retraining — you index one more vector. Which is why the evaluation is recall@k and mAP rather than accuracy, and why the embedding's geometry — not its accuracy on any fixed label set — is the thing that has to be good.
Fig. 97 — Retrieval returns an ordering, not a class, so a category that did not exist at training time costs one more indexed vector.Four modalities and what each one measures A table of four imaging modalities with the physical quantity each measures and the units it reports. CT has absolute units and can be windowed by value; MRI has none, so its intensities are only comparable within a scan. modality measures units consequence X-ray attenuation, projected arbitrary everything overlaps CT attenuation, reconstructed Hounsfield units absolute and comparable MRI proton relaxation none scanner-dependent intensities Fundus reflected light RGB colour is the finding Min-max normalising a CT throws away the property that makes it a measurement. Not normalising an the scanner in the signal. The right preprocessing is a fact about the physics, not a default in a library.
Fig. 98 — Four modalities and what each measures. The right preprocessing is a fact about the physics, not a default in a library.Fig. 99 — The ROC curve traced from strict to permissive. Every point is a threshold; a deployed system occupies exactly one of them.The same model at two prevalences One model with 90% sensitivity and 90% specificity applied to a thousand patients. At ten percent prevalence, half of its positive calls are correct. At one percent, fewer than one in ten are. Sensitivity and specificity did not change. sensitivity 90% · specificity 90% · 1000 patients prevalence 10% of the positive calls… 90 true 90 false precision 50% prevalence 1% of the positive calls… 9 true 99 false precision 8% Sensitivity and specificity are properties of the model. Precision is a property of the model and so a screening tool validated on an enriched cohort will disappoint in a clinic where the disease is rare. Ask what the prevalence was in the test set before reading any precision or F1.
Fig. 100 — The same sensitivity and specificity at two prevalences. Precision is a property of the model and the population together.Fig. 101 — Predicted probability against observed frequency. AUROC is unchanged by any monotone rescaling, so it cannot see this at all.The same model on three populations One model evaluated on an internal held-out set, on data from a different scanner at the same hospital, and on data from another institution. The internal number is the one that appears in the abstract and the external one is the one that predicts deployment. the same weights, three populations internal test 0.94 same hospital, new scanner 0.87 external hospital 0.71 Everything upstream can be right — the split, the calibration, the augmentation — and this gap opens, because every one of those checks drew from the same distribution the model was fitted on. The only measurement that answers it is data from an institution outside the training set.
Fig. 102 — One model on three populations. Every internal check drew from the distribution the model was fitted on.Fig. 103 — The whole of Book II on one plate. Two stages are architecture and three are judgement.F