Silent Failure: How Standard Metrics Overstate Capability in Deep Learning MRI Reconstruction
Abstract
Accelerated MRI reconstruction models are evaluated, and increasingly argued into clinical use, on three numbers: SSIM, PSNR and NMSE. A clinician or regulator reading a high SSIM reasonably concludes the reconstruction is faithful. We show that this conclusion does not follow, and that the gap is large enough to matter for deployment decisions. Evaluating a frozen pretrained VarNet on fastMRI with a suite that adds spectral fidelity, edge preservation, feature suppression and perceptual disagreement, we find that between 4x and 16x acceleration SSIM falls 17.5% while the reconstruction's high-frequency content, the fine detail in which small lesions live, falls 66.1%. The headline metric understates the loss by nearly a factor of four. Two results complicate the usual distribution-shift narrative: the deficit is largely present in distribution, the model discarding 40% of high-frequency content at the operating point it was trained for; and a within-dataset control shows that an out-of-distribution model apparently beating its in-distribution baseline on SSIM is an artifact of comparing SSIM across datasets rather than a property of the model. We also report a failure in our own study, because it demonstrates the reporting problem directly. Two experimental conditions silently evaluated identical data and produced a plausible, publishable table; separately, slice-level statistics made a threshold claim look supported when the effective sample size was six volumes rather than two hundred slices. Neither error is visible in a results table. We set out the reporting practices that surface such errors (volume-level uncertainty, within-dataset controls, and disclosure of the full metric suite rather than the favourable subset), and release code that regenerates every number here from the underlying records.