Chapter 7 III.7
Medical and Scientific Vision
An internal AUROC hides the two things that decide clinical usefulness: where the threshold sits and how rare the disease is.
In medicine the operating point and the prevalence decide whether a model is useful.
How this chapter is built
M3Load-bearing
The content is mathematics. Understanding is demonstrated by computation, not recall.
Five strands, not one. Mathematics is the spine; the other four are the body. A chapter cannot pay its way out of teaching with problems, nor out of problems with teaching.
Before you start
The problem
A model that reports 0.95 AUROC on held-out data can be useless in the clinic it was built for. The reasons are arithmetic — prevalence, calibration, and a shift the internal split cannot see — and each is computable before deployment.
What this chapter covers
- medical imaging
- CT
- MRI
- X-ray
- fundus imaging
- segmentation
- classification
- domain shift
- clinical evaluation
- sensitivity and specificity
- precision and recall
- F1
- AUROC
- calibration
Apparatus
The mathematics this chapter leans on, held in Book 0 so it can be assumed here without being taught here. Not a gate — follow a link when a step stops making sense.
Bayes' rule 0.PR.04 · Estimators, bias and variance 0.ST.01 · Confidence intervals 0.ST.02 · The bootstrap 0.ST.05
Notation
- 𝔼Expectation
- VarVariance
- θ̂An estimate, as against the quantity it estimates
- ℋEntropy, in nats unless bits are named
Propositions
- Prop. 1A medical image is a measurement, and what it measures decides the pipeline.CT reports absolute attenuation in calibrated units; MRI reports relaxation in arbitrary ones. The right preprocessing is a fact about the physics, not a default in a library.
- Prop. 2A model does not have a sensitivity — a threshold does.The score is continuous and the decision is binary, so every operating point trades misses against false alarms. ROC shows the whole trade; a deployed system occupies exactly one point on it.
- Prop. 3Precision is a property of the model and the population together.Hold sensitivity and specificity fixed, change how rare the disease is, and precision moves dramatically. A model validated on an enriched cohort will disappoint in a clinic.
- Prop. 4Ranking well and being right about probability are different claims.A model can order every patient correctly and still say 0.9 when it means 0.6. AUROC is invariant to any monotone rescaling of the scores, so it cannot see the difference at all.
- Prop. 5Domain shift is the failure that survives every internal check.Split by patient, calibrate, choose the augmentation carefully — and the number can still collapse at another hospital, because every one of those checks drew from the distribution the model was fitted on.
Worked problems
0/5 problems0/4 variants0/10 exercisesowes 15 more
Not yet written. At M3 this chapter owes 5 worked problems across 4 distinct variants, and 10 exercises, every one with a published solution. The build enforces that from the day the chapter is marked published.