Md. Asif Uddin

Proposition 3III.7.P0373 of 76 in the corpus

Precision is a property of the model and the population together.

Hold sensitivity and specificity fixed, change how rare the disease is, and precision moves dramatically. A model validated on an enriched cohort will disappoint in a clinic.

The same model at two prevalencesOne model with 90% sensitivity and 90% specificity applied to a thousand patients. At ten percent prevalence, half of its positive calls are correct. At one percent, fewer than one in ten are. Sensitivity and specificity did not change.sensitivity 90% · specificity 90% · 1000 patientsprevalence 10%of the positive calls…90 true90 falseprecision 50%prevalence 1%of the positive calls…9 true99 falseprecision 8%Sensitivity and specificity are properties of the model. Precision is a property of the model andso a screening tool validated on an enriched cohort will disappoint in a clinic where the disease is rare.Ask what the prevalence was in the test set before reading any precision or F1.
Fig. 3 — The same sensitivity and specificity at two prevalences. Precision is a property of the model and the population together.

Demonstration

Take a model at 90% sensitivity and 90% specificity — respectable numbers — and apply it to a thousand patients.

At 10% prevalence. 100 diseased, 900 healthy. It finds 90 of the 100 and falsely flags 90 of the 900. Of 180 positive calls, half are right. Precision 50%.

At 1% prevalence. 10 diseased, 990 healthy. It finds 9 and falsely flags 99. Of 108 positive calls, 9 are right. Precision 8%.

The model did not change. Sensitivity and specificity did not change. What changed is the population, and precision moved by a factor of six.

This is why a research dataset assembled with a comfortable fraction of positives — because that is what makes training tractable — cannot support a claim about a screening clinic where prevalence is a hundredth of that. The paper’s precision, F1 and PPV are all reported on the wrong denominator.

It also explains a practice that looks like sleight of hand until you see the arithmetic: quoting specificity at a fixed high sensitivity. At 1% prevalence, moving specificity from 90% to 99% takes false positives from 99 to about 10 and precision from 8% to nearly half. In the rare-disease regime, specificity is where the usefulness lives.

Corollary

Before reading any precision or F1, find the prevalence in the test set and compare it with the deployment setting. If they differ by an order of magnitude, recompute — the sensitivity and specificity transfer, and everything derived from them does not.

Sources