Reference-Free Per-Mask Reliability Assessment for Whole-Body Anatomical Segmentation Across CT and MRI
Abstract
Segmentation tools such as TotalSegmentator increasingly serve as black-box imaging tools inside automated clinical AI pipelines, yet return masks with no per-output reliability signal. On 57 held-out CT subjects, 25.5\% of masks fall below an IoU of 0.90 with per-structure pass rates spanning 2.7\% to 100\%; on an in-sample MRI subset, 77.0\% fall below the same bar. The failures are largely silent: of masks failing on overlap, 52.8\% (CT) and 69.5\% (MRI) pass a naive volume check. We present a reference-free reliability layer that predicts per-mask quality from nineteen geometric and intensity descriptors computed from the predicted mask and source image alone, with no reference annotation, auxiliary network, or segmenter access. It reaches ROC-AUC 0.871 (95\% CI 0.852--0.892) on held-out CT and 0.919 (0.890--0.947) on the in-sample MRI subset, improving over organ-identity baselines by 0.110 and 0.104 AUC points, respectively, and transfers to anatomical families withheld from training (mean AUC 0.779). We show the fixed IoU threshold encodes a size-dependent physical standard and replace classification with IoU regression so the operating point is chosen at deployment. Scores are calibrated where data permit (ECE 0.029 to 0.013 on CT) and equipped with conformal intervals (coverage 52.0\% to 80.7\% on CT, 45.4\% to 78.6\% on MRI, at nominal 80\%), so reported confidence can be acted upon.