Plate III
Automated Renal Reporting
Measure First, Write Second: Verified Generation of the Kidney Section of a Radiology Report from Abdominal CT
The problem
Every report-generation paper is challenged on the same thing: the model writes a fluent sentence containing a number nobody measured. Most answer with better metrics. An image-to-text model trained to produce radiology prose has no mechanism preventing an invented value — only a distribution that makes one less likely.
Approach
Measure first, write second. The pipeline does not predict text from images. It segments structures, computes numbers from those structures, fills the numbers into a template, and only then lets a language model rephrase. Generation happens last and cannot introduce a quantity, because every quantity is already fixed before it runs.
The verifier enforces it. If the rephrased sentence drops a measured value, the output is rejected and the template version is used instead. Hallucination becomes structurally impossible rather than statistically unlikely — an architectural answer to the objection, not a metric.
The phase gate bounds the claims. A classifier decides native against contrast. On a native study the report cannot mention enhancement, because enhancement is not assessable without contrast. On a contrast study calculus detection is suppressed, because opacified urine occupies the same attenuation range as stone. The report says what was assessable rather than answering regardless of input — the restraint a radiologist already shows.
The sinus region substitutes for annotation. The collecting system is what dilates in hydronephrosis and where stones sit, and no public dataset segments it. So it is derived geometrically: the convex hull of the kidney minus the kidney mask isolates the sinus, then a density split separates fat from fluid. Fluid volume over kidney volume is the feature. No labels required, and no radiologist to ask for them.
What learns and what does not
Three trained components — kidney segmentation, lesion segmentation, and the hydronephrosis classifier. Everything else is deterministic: measurement is geometry on masks, calculus detection is a threshold, template filling is string substitution, the verifier is a check.
That split is the point. A deterministic component can be inspected, debugged and explained, and it cannot drift.
Evaluation
Merlin’s own patient-level split is preserved to prevent leakage. Segmentation against held-out human masks by Dice. Measurement is run twice, on trusted masks and on predicted ones, which separates segmentation error from measurement error rather than reporting their sum. Findings by sensitivity, PPV and AUC; reports by BLEU and RadGraph F1 against real radiologist text, alongside the verifier’s pass rate.
Two ablations carry the argument: whether the derived sinus feature beats intensity thresholding for hydronephrosis, and whether the verifier measurably reduces value omission against unconstrained rephrasing.
External validation
Everything above is trained and tuned on public data. The system is then run, unchanged, against 1,062 institutional KUB CT studies with paired radiologist reports — de-identified, each linking a volumetric scan to its DICOM-derived metadata and the report a radiologist actually wrote, with multi-label annotations for urolithiasis, hydronephrosis, renal cysts, pyelonephritis, cystitis, nephrocalcinosis, renal masses and normal studies.
The cohort is deliberately stratified toward the cases where renal measurement is hardest and matters most: hydronephrosis, staghorn calculi, and atrophic kidneys. A held-out set drawn to be easy would flatter the method rather than test it.
That cohort cannot be released, and the split is what makes the restriction a design choice rather than a limitation. A reviewer is right to distrust a model trained on data nobody else can obtain — the result cannot be checked or built on. An external validation set is the opposite case: it is the part a reader least needs to hold, because the claim it supports is that a system built entirely from public sources still works on data it has never seen, from a different scanner population, graded against a different set of radiologists.
Reproducibility
Trained and validated on public data alone, with the cohort ID list and the frozen configuration released. Every number in the main results can be reproduced from public sources; the institutional cohort adds a second, external check that cannot be reproduced and is reported separately for that reason.
Status
In preparation. Results are held until publication.