Book III · Closing
Vision, gathered
Book II is one plate. Pixels are normalised into a tensor whose axes carry the meaning, mixed either by a convolution that assumes locality or by attention over patches that assumes nothing, built into a hierarchy nobody specified, read by a head chosen for the question, and turned into a number that someone acts on.
Two of those stages are architecture and three are judgement. Most of the literature is about the second row. Most of what goes wrong in a deployed system is in the last one.
CNN or transformer
The honest comparison is not a winner but a set of conditions.
| convolution | attention over patches | |
|---|---|---|
| prior | locality, translation equivariance | none |
| data needed | works from thousands | wants a hundred million to beat the prior |
| cost in length | linear | quadratic |
| path between distant pixels | grows with depth | one step |
| resolution | any, without ceremony | fixed by patch size and position embeddings |
| pyramid | free | has to be rebuilt, as Swin does |
The field’s own corrections point one way. Swin reintroduced windows and a hierarchy. ConvNeXt matched a transformer by modernising a ResNet’s recipe. Swin UNETR V2 had to insert a convolution at the head of every encoder stage. Priors are a constraint whose value depends on how much data you have, and medical datasets are small by construction.
Supervised or self-supervised
| supervised | self-supervised | |
|---|---|---|
| what it costs | annotation | compute and unlabelled data |
| what it learns | features useful for one label set | features useful for many |
| frozen quality | good on the source task | good in general, for joint-embedding methods |
| where it wins | plentiful labels, fixed task | scarce labels, unknown downstream task |
The reason this chapter sits in a medical book rather than a general one is that the binding constraint in medical imaging is a specialist’s time, not storage. Images are comparatively abundant and labels are not, which is exactly the regime self-supervision was built for.
How to compare two vision models
Freeze both. Train a linear probe on the same few hundred labelled images from your own domain, with the same preprocessing and the same splits, over several seeds. Report the spread, not the best run.
Then, and only then, fine-tune. If fine-tuning beats the probe by a large margin on a small dataset, be suspicious before being pleased — that gap is what memorisation looks like from the outside.
The equations
An image
x ∈ ℝ^(H × W × C)Rows, columns, channels. The numbers are only numbers; the axes are what make it a picture.
Normalisation
x ← (x − μ) / σ, μ and σ from the training setRecomputing them at test time leaks between test examples, and the leak is silent.
Convolution output size
out = ⌊(in + 2·pad − k) / stride⌋ + 1Nothing else decides it. Most shape errors are this formula applied by the framework and not by the author.
Receptive field, stride-1 3×3 stack
r = 2L + 1 after L layersLinear in depth, and the effective field is much smaller than the nominal one.
Patches
n = (H/p)(W/p), attention cost ∝ n²Halving the patch quadruples the tokens and multiplies the bill by sixteen.
Sensitivity and specificity
TP/(TP+FN) and TN/(TN+FP)Properties of a threshold, not of a model. Lowering it raises one and lowers the other, always.
Precision
TP/(TP+FP)Depends on prevalence. The same model at 10% and 1% prevalence reports 50% and 8%.
Dice
2|A ∩ B| / (|A| + |B|)Region overlap. High while a small lesion is missed entirely, low on a well-found structure with a soft edge.
The vocabulary
- Channel
- One measurement per pixel. Colour for a photograph, a sequence for MRI, a modality for anything multispectral.
- Feature map
- The output of one filter at every position: how strongly that pattern was present, and where.
- Receptive field
- The input pixels that can influence one unit. Computable from the architecture before any training.
- Output stride
- How many input pixels one output position covers. It is the finest structure the head can localise.
- Inductive bias
- An assumption welded into the architecture rather than learned. Locality and translation equivariance are the two that matter here.
- Patch
- A fixed square of pixels treated as one token. The /16 in ViT-B/16.
- Pretext task
- A problem whose answer you know because you created the corruption. Nobody wants its output; they want what solving it required.
- Collapse
- Every input mapping to one vector. Not an instability — the global optimum of an objective that only rewards agreement.
- Linear probe
- A logistic regression on frozen features. The standard measurement of a representation, and the standard way to compare two encoders.
- Domain shift
- The deployment distribution differing from the training one. The failure that survives every check drawn from the training distribution.
- Calibration
- Whether a stated probability matches an observed frequency. Independent of ranking, and invisible to AUROC.
- Prevalence
- How common the finding is in the population being tested. It moves precision and leaves sensitivity alone.
Where it is usually got wrong
said Transformers replaced convolutions in vision.
in fact ViT's own ablation is a crossover with a threshold near a hundred million images. Below it convolution wins, and Swin, ConvNeXt and Swin UNETR V2 all put a locality prior back. Match the prior to the data budget.
said A bigger input resolution is always better.
in fact It is always more expensive, and for a ViT it costs quadratically. Whether it helps depends on whether the structure you care about was smaller than a patch — which is a question about your data, not about the model.
said Self-supervised learning needs no labels, so it is free.
in fact It needs no annotations and a great deal of compute and data. What it does not need is a radiologist, which is why it matters in medicine specifically.
said The augmentation pipeline is a detail.
in fact It is the specification of what the representation should ignore. SimCLR's ablation found it mattered more than the architecture, and colour jitter on a fundus photograph tells the model to ignore the finding.
said Pooling and stride throw away redundant information.
in fact They throw away position, irreversibly. Every segmentation architecture is shaped by the need to get it back, which is what the skip connections in a U-Net are for.
said A high AUROC means the model is ready.
in fact AUROC is a statement about ranking at every threshold at once. It says nothing about where you will operate, nothing about precision at your prevalence, and nothing about calibration.
said Dice is the segmentation metric.
in fact It measures overlap. A clinician usually wants to know whether the lesion was found. Report per-lesion detection beside it.
said Our model generalises — we held out a test set.
in fact Held out from the same distribution. External validation on data from an institution that contributed nothing to training is the only measurement that answers the question.
said Pretrained on ImageNet, so it understands images.
in fact Its early layers are edge detectors, which transfer almost anywhere. Its late layers are about the thousand ImageNet classes, and transfer to a retinal photograph considerably less well.
The problem set
Cross-chapter problems, which is where the mathematics of a book stops being sectional. Not yet written for this book.
Book III problem set →Laboratory
Things to go and run. A worked problem is checked against arithmetic; these are checked against a machine.
Take one image, shuffle its pixels with a fixed permutation, and train the same MLP on the original and the shuffled dataset. Then do it with a small CNN.
Chapter I. The MLP should be indifferent and the CNN should not. That gap is the locality prior, measured.
Compute the receptive field of a backbone you use, by hand, layer by layer. Then check it against the size of the structure your task depends on.
Chapter II. If it does not cover it, no amount of data will help, and you now know that before training.
Train a linear probe on frozen features from three backbones — a supervised CNN, DINOv2, and a masked autoencoder — on a few hundred labelled images from your own domain.
Chapter III and V. This is the cheapest honest comparison of representations available, and it takes an afternoon.
Take a ViT pretrained at 224 and evaluate it at 448, once with interpolated positional embeddings and once without.
Chapter IV. The version without is the commonest silent bug in this literature.
Write out what each transform in your augmentation pipeline asserts about your domain, then check the list with someone who reads the images clinically.
Chapter V. Ten minutes, and it is the difference between ignoring nuisance and ignoring the diagnosis.
Take a trained classifier and plot its reliability diagram. Then fit a temperature on held-out data and plot it again.
Chapter VII. Accuracy and AUROC will be identical before and after, which is the point.
Recompute the precision of a published medical model at the prevalence of the clinic it is meant for, using its reported sensitivity and specificity.
Chapter VII. Do this with three papers and the habit will be permanent.
A project, end to end. Pick a public medical dataset with site or scanner metadata. Split it by patient and train. Then split it by site, retrain, and compare. Report Dice, per-lesion detection, calibration error and both splits' numbers.
The gap between the two splits is the result. Everything else is method.
Read these
- Gradient-Based Learning Applied to Document Recognition (LeNet)Where the shape comes from, and worth reading for how much of it survived.
- ImageNet Classification with Deep CNNs (AlexNet)The moment learned features stopped being controversial.
- Deep Residual Learning for Image RecognitionThe one that settled depth. Everything that stacks deeply inherits it, transformers included.
- U-Net: Convolutional Networks for Biomedical Image SegmentationStill the shape of nearly every dense prediction network, and the skips are the reason.
- An Image is Worth 16x16 WordsRead the dataset-size ablation rather than the headline table. The ablation is where the science is.
- Swin TransformerA transformer with the pyramid and the locality put back, and honest about why.
- A Simple Framework for Contrastive Learning (SimCLR)The augmentation ablation is the most useful table in the self-supervised literature.
- Emerging Properties in Self-Supervised Vision Transformers (DINO)Attention maps that segment objects nobody labelled.
- Masked Autoencoders Are Scalable Vision LearnersWhy the mask ratio has to be 75%, which is a fact about images rather than about the method.
- nnU-NetThe argument that on 3D medical segmentation the architecture was never the bottleneck.
- Variable Generalization Performance of a Deep Learning Model to Detect PneumoniaThe models learned which hospital the radiograph came from. Read this before claiming generalisation.
- On Calibration of Modern Neural NetworksNetworks got more accurate and less calibrated at the same time.
What comes next
Book III keeps the transformer and changes what the tokens are made of, and the two meet again in Book IV: a vision-language model is this book's encoder and that book's decoder, joined by cross-attention, and the whole difficulty is making two representations comparable. What carries forward is not the architecture — it is the last chapter. Every claim about a model is a claim about a measurement, and the measurement is only as good as the population it was taken on.