Does Confidence Track Visibility? Diagnosing Blind Confidence in Cloud-Occluded Flood Segmentation
Abstract
Optical flood-mapping models are least reliable exactly when they are most needed: heavy cloud cover coincides with the monsoon and cyclone seasons that produce the floods these models must detect. We ask a narrow, falsifiable question – does a model’s predicted confidence track the real information loss from cloud occlusion, or does it stay artificially high as actual skill collapses? We build a controlled cloud-severity ladder over real Sentinel-1/Sentinel-2 chips from Sen1Floods11 and evaluate four flood classifiers – an optical U-Net, a cloud-invariant SAR U-Net, an early-fusion SAR+optical U-Net, and a classical NDWI rule – across seven severities. The optical model’s confidence-minus-accuracy gap (GapAUC) is significantly positive (+0.047, 95% CI [+0.027, +0.070], n = 64), while SAR’s gap is an order of magnitude smaller, exactly as its cloud-invariant design predicts. Water-class IoU shows this understates the damage: raw accuracy suggests mild degradation, while IoU collapses from 0.457 to 0.146. A false-negative/false-positive decomposition shows the fusion model’s apparently good calibration masks a different failure – a sharp rise in false alarms, not missed detections, under heavy cloud. Finally, a per-severity selective-prediction layer targeting a 10% missed-flood rate certifies no threshold at any severity for any model, a result we attribute to insufficient model accuracy at pilot scale, not the calibration procedure. Code and configuration: https://anonymous.4open.science/r/blind-confidence-flood-mapping-9123/