Md. Asif Uddin

    Book III · Closing

    Vision, gathered

    7 chapters · 26 propositions

    Book II is one plate. Pixels are normalised into a tensor whose axes carry the meaning, mixed either by a convolution that assumes locality or by attention over patches that assumes nothing, built into a hierarchy nobody specified, read by a head chosen for the question, and turned into a number that someone acts on.

    Two of those stages are architecture and three are judgement. Most of the literature is about the second row. Most of what goes wrong in a deployed system is in the last one.

    CNN or transformer

    The honest comparison is not a winner but a set of conditions.

    convolutionattention over patches
    priorlocality, translation equivariancenone
    data neededworks from thousandswants a hundred million to beat the prior
    cost in lengthlinearquadratic
    path between distant pixelsgrows with depthone step
    resolutionany, without ceremonyfixed by patch size and position embeddings
    pyramidfreehas to be rebuilt, as Swin does

    The field’s own corrections point one way. Swin reintroduced windows and a hierarchy. ConvNeXt matched a transformer by modernising a ResNet’s recipe. Swin UNETR V2 had to insert a convolution at the head of every encoder stage. Priors are a constraint whose value depends on how much data you have, and medical datasets are small by construction.

    Supervised or self-supervised

    supervisedself-supervised
    what it costsannotationcompute and unlabelled data
    what it learnsfeatures useful for one label setfeatures useful for many
    frozen qualitygood on the source taskgood in general, for joint-embedding methods
    where it winsplentiful labels, fixed taskscarce labels, unknown downstream task

    The reason this chapter sits in a medical book rather than a general one is that the binding constraint in medical imaging is a specialist’s time, not storage. Images are comparatively abundant and labels are not, which is exactly the regime self-supervision was built for.

    How to compare two vision models

    Freeze both. Train a linear probe on the same few hundred labelled images from your own domain, with the same preprocessing and the same splits, over several seeds. Report the spread, not the best run.

    Then, and only then, fine-tune. If fine-tuning beats the probe by a large margin on a small dataset, be suspicious before being pleased — that gap is what memorisation looks like from the outside.

    The whole of Book II on one plateA vertical map from raw pixels to a decision. Pixels are normalised, mixed either by convolution or by attention over patches, built into a hierarchy of features, read by a task head, and finally turned into a number that someone acts on. Each stage is labelled with the chapter that covers it.from pixels to a decisionpixels, channels, normalisationCh. Iconvolution or patchesCh. II, IVa learned hierarchy of featuresCh. III, Va head for the question askedCh. VIa number somebody acts onCh. VIIone imageTwo stages are architecture and three are judgement. Most of the literature is about the second row;goes wrong in a deployed system is in the last one.If a stage here is not one you could explain to somebody else, the chapter is named beside it.
    Fig. — — The whole of Book II on one plate. Two stages are architecture and three are judgement.

    The equations

    1. An image

      x ∈ ℝ^(H × W × C)

      Rows, columns, channels. The numbers are only numbers; the axes are what make it a picture.

    2. Normalisation

      x ← (x − μ) / σ,   μ and σ from the training set

      Recomputing them at test time leaks between test examples, and the leak is silent.

    3. Convolution output size

      out = ⌊(in + 2·pad − k) / stride⌋ + 1

      Nothing else decides it. Most shape errors are this formula applied by the framework and not by the author.

    4. Receptive field, stride-1 3×3 stack

      r = 2L + 1  after L layers

      Linear in depth, and the effective field is much smaller than the nominal one.

    5. Patches

      n = (H/p)(W/p),   attention cost ∝ n²

      Halving the patch quadruples the tokens and multiplies the bill by sixteen.

    6. Sensitivity and specificity

      TP/(TP+FN)   and   TN/(TN+FP)

      Properties of a threshold, not of a model. Lowering it raises one and lowers the other, always.

    7. Precision

      TP/(TP+FP)

      Depends on prevalence. The same model at 10% and 1% prevalence reports 50% and 8%.

    8. Dice

      2|A ∩ B| / (|A| + |B|)

      Region overlap. High while a small lesion is missed entirely, low on a well-found structure with a soft edge.

    The vocabulary

    Channel
    One measurement per pixel. Colour for a photograph, a sequence for MRI, a modality for anything multispectral.
    Feature map
    The output of one filter at every position: how strongly that pattern was present, and where.
    Receptive field
    The input pixels that can influence one unit. Computable from the architecture before any training.
    Output stride
    How many input pixels one output position covers. It is the finest structure the head can localise.
    Inductive bias
    An assumption welded into the architecture rather than learned. Locality and translation equivariance are the two that matter here.
    Patch
    A fixed square of pixels treated as one token. The /16 in ViT-B/16.
    Pretext task
    A problem whose answer you know because you created the corruption. Nobody wants its output; they want what solving it required.
    Collapse
    Every input mapping to one vector. Not an instability — the global optimum of an objective that only rewards agreement.
    Linear probe
    A logistic regression on frozen features. The standard measurement of a representation, and the standard way to compare two encoders.
    Domain shift
    The deployment distribution differing from the training one. The failure that survives every check drawn from the training distribution.
    Calibration
    Whether a stated probability matches an observed frequency. Independent of ranking, and invisible to AUROC.
    Prevalence
    How common the finding is in the population being tested. It moves precision and leaves sensitivity alone.

    Where it is usually got wrong

    • said Transformers replaced convolutions in vision.

      in fact ViT's own ablation is a crossover with a threshold near a hundred million images. Below it convolution wins, and Swin, ConvNeXt and Swin UNETR V2 all put a locality prior back. Match the prior to the data budget.

    • said A bigger input resolution is always better.

      in fact It is always more expensive, and for a ViT it costs quadratically. Whether it helps depends on whether the structure you care about was smaller than a patch — which is a question about your data, not about the model.

    • said Self-supervised learning needs no labels, so it is free.

      in fact It needs no annotations and a great deal of compute and data. What it does not need is a radiologist, which is why it matters in medicine specifically.

    • said The augmentation pipeline is a detail.

      in fact It is the specification of what the representation should ignore. SimCLR's ablation found it mattered more than the architecture, and colour jitter on a fundus photograph tells the model to ignore the finding.

    • said Pooling and stride throw away redundant information.

      in fact They throw away position, irreversibly. Every segmentation architecture is shaped by the need to get it back, which is what the skip connections in a U-Net are for.

    • said A high AUROC means the model is ready.

      in fact AUROC is a statement about ranking at every threshold at once. It says nothing about where you will operate, nothing about precision at your prevalence, and nothing about calibration.

    • said Dice is the segmentation metric.

      in fact It measures overlap. A clinician usually wants to know whether the lesion was found. Report per-lesion detection beside it.

    • said Our model generalises — we held out a test set.

      in fact Held out from the same distribution. External validation on data from an institution that contributed nothing to training is the only measurement that answers the question.

    • said Pretrained on ImageNet, so it understands images.

      in fact Its early layers are edge detectors, which transfer almost anywhere. Its late layers are about the thousand ImageNet classes, and transfer to a retinal photograph considerably less well.

    The problem set

    Cross-chapter problems, which is where the mathematics of a book stops being sectional. Not yet written for this book.

    Book III problem set →

    Laboratory

    Things to go and run. A worked problem is checked against arithmetic; these are checked against a machine.

    1. Take one image, shuffle its pixels with a fixed permutation, and train the same MLP on the original and the shuffled dataset. Then do it with a small CNN.

      Chapter I. The MLP should be indifferent and the CNN should not. That gap is the locality prior, measured.

    2. Compute the receptive field of a backbone you use, by hand, layer by layer. Then check it against the size of the structure your task depends on.

      Chapter II. If it does not cover it, no amount of data will help, and you now know that before training.

    3. Train a linear probe on frozen features from three backbones — a supervised CNN, DINOv2, and a masked autoencoder — on a few hundred labelled images from your own domain.

      Chapter III and V. This is the cheapest honest comparison of representations available, and it takes an afternoon.

    4. Take a ViT pretrained at 224 and evaluate it at 448, once with interpolated positional embeddings and once without.

      Chapter IV. The version without is the commonest silent bug in this literature.

    5. Write out what each transform in your augmentation pipeline asserts about your domain, then check the list with someone who reads the images clinically.

      Chapter V. Ten minutes, and it is the difference between ignoring nuisance and ignoring the diagnosis.

    6. Take a trained classifier and plot its reliability diagram. Then fit a temperature on held-out data and plot it again.

      Chapter VII. Accuracy and AUROC will be identical before and after, which is the point.

    7. Recompute the precision of a published medical model at the prevalence of the clinic it is meant for, using its reported sensitivity and specificity.

      Chapter VII. Do this with three papers and the habit will be permanent.

    8. A project, end to end. Pick a public medical dataset with site or scanner metadata. Split it by patient and train. Then split it by site, retrain, and compare. Report Dice, per-lesion detection, calibration error and both splits' numbers.

      The gap between the two splits is the result. Everything else is method.

    Read these

    What comes next

    Book III keeps the transformer and changes what the tokens are made of, and the two meet again in Book IV: a vision-language model is this book's encoder and that book's decoder, joined by cross-attention, and the whole difficulty is making two representations comparable. What carries forward is not the architecture — it is the last chapter. Every claim about a model is a claim about a measurement, and the measurement is only as good as the population it was taken on.