Projects12 of 15
MONAI
The metric that hid the failures
Where it was found
MONAI is the medical imaging library that most segmentation work in this field is built on. I ran its segmentation metrics against the degenerate cases retinal lesion work produces: empty masks, single-pixel structures, and predictions that miss entirely. That is where a widely used library is least tested, and where my own data lives.
The bug
HausdorffDistanceMetric answers the same input two different ways. For a prediction that found nothing:
- the maximum returns
inf, which is honest; - every percentile returns
nan.
HD95 is the headline metric in medical segmentation, and nan is that metric’s not-applicable sentinel, so do_metric_reduction drops it before averaging. The failures leave the score.
What it does to a result
Take a hundred images where the only variable is how many the model misses. The reported HD95 does not move. A model that finds nothing in ninety-nine of them scores the same as one that finds the lesion every time, because the average was taken over the one case that worked.
The cause
torch.quantile interpolates between two infinities, inf + (inf - inf) * frac, which gives nan. The maximum and minimum paths escape because they do not interpolate.
It reproduces on the released 1.6.0 as well as on the development branch, so it affects installs in use, not only unreleased code.
The fix
Return the infinity directly. Twelve regression tests cover both entry points at every percentile. Without the source change, eight fail and four pass, and the four are exactly the paths that do not interpolate: the fix checking its own diagnosis.
Status
Reported as issue #9095. The fix, pull request #9096, was reviewed and merged into MONAI on 6 October 2026, closing the issue.