Right Number, Wrong Conclusion: Seven Failures in Clinical ML Evaluation
Saanvi S Subramanian
Abstract
Biomedical machine-learning results may support conclusions of larger scope than the analyses justify even when the reported numbers are correct. We examine this problem through a self-audit of one EEG-guided lesion-localization project. We reconstruct seven claims that our analyses initially overstated and identify the check that exposed each problem. The failures involved an inappropriate evaluation baseline, a missing random-initialized control, an invalid ablation, unmatched perturbations, perturbation confounds, a seed-sensitive threshold, and an incorrectly described cohort. The corrections changed both numerical results and scientific interpretation. A spatially blind predictor reached $1.24$--$1.37\times$ density-matched chance from class-frequency structure alone, changing a $2.00\times$-chance localization result to $1.62\times$ the relevant blind floor. An ablation that appeared to show that coordinates were unimportant still exposed sensor identity through the readout. After removing that route, coordinate corruption reduced performance from $1.53\times$ to $1.22\times$. Confound-controlled probes also showed that sensor identity had a much larger effect than temporal order, with costs of $1.338\pm0.054$ and $0.163\pm0.071$. A $90\%$-accurate spatial prior that appeared beneficial with three seeds was indistinguishable from no prior with eight unique seeds ($p=0.547$). Finally, a result described as covering 262 cases across four releases had been computed on 25 cases from one site. Reanalysis of all 136 scoreable cases preserved the qualitative conclusion within every release. Six of the seven problems had a cheap diagnostic: one additional run, control, configuration, count, or explicit check, except for the seed-sensitive threshold. These cases show that we must use stronger evaluations to determine whether the evidence supports the specific scientific claim instead of relying solely on the accuracy of a computed metric.
Chat is not available.
Successful Page Load