Silent Failures in Degraded Chest X-Rays: Triage in Resource-Constrained Settings
Abstract
Pneumonia remains one of the leading killers of children under five, with Sub-Saharan Africa carrying a disproportionate share of that burden. Chest X-rays are the standard diagnostic tool, but the region's radiologist shortage means most X-rays never get read in time. AI assisted triage seems like the fix. Nobody has really asked what happens to that AI when the imaging hardware is bad. Older X-ray generators, worn sensors, and low-bandwidth transmission all distort inputs in ways that can cause a classifier to fail silently: producing a confident, wrong answer instead of flagging the case for review. We ask which combination of lightweight backbone and uncertainty method keeps failure detection reliable as image quality degrades. We fine-tune three compact CNN backbones on PneumoniaMNIST: MobileNetV3-Small, EfficientNet-B0, and ConvNeXt-Tiny. Each one is paired with four uncertainty quantification methods: softmax confidence, MC Dropout, temperature scaling, and ensemble predictive entropy. In total we run 650 combinations of backbone, method, corruption type, and severity level, across the 13 MedMNIST-C corruptions (noise, blur, brightness, contrast, gamma, JPEG compression, pixelation) at five severity levels. For each one we measure accuracy, calibration (Expected Calibration Error), and failure detection AUROC. All three backbones start out close, with clean accuracy between 83% and 86%, but they diverge sharply once corruption is introduced. EfficientNet-B0 drops to about 71% accuracy at the mildest corruption level, and selective prediction can't save it: at severity 1, coverage at a 90% accuracy target falls to 0.0 for almost every corruption type. MobileNetV3-Small and ConvNeXt-Tiny stay within two points of their clean baseline under the same conditions. ConvNeXt-Tiny gets the best AUROC (0.961), but its calibration is the worst of the three (ECE 16.5%), and temperature scaling doesn't fix it. Ensemble entropy is the strongest at catching failures (FD-AUROC 0.804 clean), while temperature scaling gives the lowest ECE at no extra inference cost. On the clean test set (n = 624), the best results per backbone all come from temperature scaling except the ensemble: MobileNetV3-Small reaches 85.1% accuracy, 0.928 AUROC, 10.7% ECE, and 0.775 FD-AUROC. EfficientNet-B0 reaches 85.1% accuracy, 0.915 AUROC, 11.0% ECE, and 0.778 FD-AUROC. ConvNeXt-Tiny reaches 83.3% accuracy, 0.961 AUROC, 16.5% ECE, and 0.806 FD-AUROC. The ensemble, using entropy, reaches 85.6% accuracy, 0.949 AUROC, 12.8% ECE, and 0.804 FD-AUROC. MobileNetV3-Small paired with temperature scaling is the practical choice: 2.5M parameters, clean accuracy on par with EfficientNet-B0, and the lowest ECE of the three backbones. In a simulated clinic receiving 100 chest X-rays a day under mild corruption, deferring the most uncertain 20-25% of cases to a human reviewer holds 90% accuracy on the rest, so roughly 75 patients can be auto-triaged and only 25 need specialist attention. Past moderate corruption severity, uncertainty signals lose discriminative power, and periodic image quality checks become necessary regardless of backbone or method. Calibration, not raw accuracy, determines whether a lightweight classifier is safe to deploy in a hardware-constrained clinical setting.