Md. Asif Uddin

    Plate III

    Automated Renal Reporting

    Measure First, Write Second: Verified Generation of the Kidney Section of a Radiology Report from Abdominal CT

    In preparation
    Year
    2026
    Supervisor
    Niloy Farhan
    Datasets
    KiTS23 and AMOS22 — ~990 human-annotated cases for kidney segmentation · KiTS23 — lesion segmentation · Merlin — 787 report-derived hydronephrosis labels; patient-level split 7,466 / 2,451 / 2,504 · 1,062 institutional KUB CT studies with paired radiologist reports — external validation only, never trained on

    The problem

    Every report-generation paper is challenged on the same thing: the model writes a fluent sentence containing a number nobody measured. Most answer with better metrics. An image-to-text model trained to produce radiology prose has no mechanism preventing an invented value — only a distribution that makes one less likely.

    Approach

    Measure first, write second. The pipeline does not predict text from images. It segments structures, computes numbers from those structures, fills the numbers into a template, and only then lets a language model rephrase. Generation happens last and cannot introduce a quantity, because every quantity is already fixed before it runs.

    The verifier enforces it. If the rephrased sentence drops a measured value, the output is rejected and the template version is used instead. Hallucination becomes structurally impossible rather than statistically unlikely — an architectural answer to the objection, not a metric.

    The phase gate bounds the claims. A classifier decides native against contrast. On a native study the report cannot mention enhancement, because enhancement is not assessable without contrast. On a contrast study calculus detection is suppressed, because opacified urine occupies the same attenuation range as stone. The report says what was assessable rather than answering regardless of input — the restraint a radiologist already shows.

    The sinus region substitutes for annotation. The collecting system is what dilates in hydronephrosis and where stones sit, and no public dataset segments it. So it is derived geometrically: the convex hull of the kidney minus the kidney mask isolates the sinus, then a density split separates fat from fluid. Fluid volume over kidney volume is the feature. No labels required, and no radiologist to ask for them.

    What learns and what does not

    Three trained components — kidney segmentation, lesion segmentation, and the hydronephrosis classifier. Everything else is deterministic: measurement is geometry on masks, calculus detection is a threshold, template filling is string substitution, the verifier is a check.

    That split is the point. A deterministic component can be inspected, debugged and explained, and it cannot drift.

    Evaluation

    Merlin’s own patient-level split is preserved to prevent leakage. Segmentation against held-out human masks by Dice. Measurement is run twice, on trusted masks and on predicted ones, which separates segmentation error from measurement error rather than reporting their sum. Findings by sensitivity, PPV and AUC; reports by BLEU and RadGraph F1 against real radiologist text, alongside the verifier’s pass rate.

    Two ablations carry the argument: whether the derived sinus feature beats intensity thresholding for hydronephrosis, and whether the verifier measurably reduces value omission against unconstrained rephrasing.

    External validation

    Everything above is trained and tuned on public data. The system is then run, unchanged, against 1,062 institutional KUB CT studies with paired radiologist reports — de-identified, each linking a volumetric scan to its DICOM-derived metadata and the report a radiologist actually wrote, with multi-label annotations for urolithiasis, hydronephrosis, renal cysts, pyelonephritis, cystitis, nephrocalcinosis, renal masses and normal studies.

    The cohort is deliberately stratified toward the cases where renal measurement is hardest and matters most: hydronephrosis, staghorn calculi, and atrophic kidneys. A held-out set drawn to be easy would flatter the method rather than test it.

    That cohort cannot be released, and the split is what makes the restriction a design choice rather than a limitation. A reviewer is right to distrust a model trained on data nobody else can obtain — the result cannot be checked or built on. An external validation set is the opposite case: it is the part a reader least needs to hold, because the claim it supports is that a system built entirely from public sources still works on data it has never seen, from a different scanner population, graded against a different set of radiologists.

    Reproducibility

    Trained and validated on public data alone, with the cohort ID list and the frozen configuration released. Every number in the main results can be reproduced from public sources; the institutional cohort adds a second, external check that cannot be reproduced and is reported separately for that reason.

    Status

    In preparation. Results are held until publication.