Trivial baselines catch label leakage that cross-validation misses: a flood-mapping case study on a data-scarce West African coast
Jerry E Kudaya
Abstract
On coastlines without gauges or weather stations, machine-learning flood models are often trained and evaluated against the same remotely sensed proxy (e.g., a Sentinel-1 backscatter-change flood mask), leaving no independent check that the proxy measures flooding. We document a case in which this failure remained undetected. A distance-to-water predictor, produced by an inverted distance transform, encoded the flood label on 61% of samples and accounted for 83% of model importance. The resulting susceptibility model achieved 81% accuracy under strict chronological validation, varied by only 0.03 in ROC-AUC across four classifier families (logistic regression, random forest, gradient boosting, and LightGBM), and calibrated to a Brier skill score of 0.33. All of these checks suggested robust performance. However, a single threshold on the leaked feature outperformed the 28-feature model ($\kappa$ 0.503 vs 0.383), and the model's land-only ROC-AUC was 0.515. Reformulating the target to a quantity the observing system can measure recovers substantive skill ($R^2$ 0.487). We argue that trivial baselines, reported alongside every model, provide a cheap and transferable safeguard that leakage audits and multi-model benchmarking may not.
Chat is not available.
Successful Page Load