Which Floods Should We Label? Environment-Aware Event Selection for Data-Efficient Flood Mapping
Abstract
Expert flood labels are commissioned one disaster at a time, so the choice of which events to annotate is an upstream data-curation decision. Using the 11 hand-labeled events of Sen1Floods11 with a fixed segmentation protocol and leave-one-event-out evaluation, we first measure a large cross-event transfer gap of 0.198 water IoU and show that most of the remaining variation is set by intrinsic target-event difficulty, which bounds what any selection rule can achieve. Within that bound, we represent each candidate event by environmental descriptors computable before annotation and select subsets by core-set coverage. At tight budgets this improves the cost–accuracy frontier: +0.047 water IoU over continental stratification, with 10 of 11 held-out events improved (permutation p = 0.032), while matching the strongest accuracy baseline with 45% fewer labeled chips. Under matched chip budgets the advantage remains positive but is statistically unresolved at eleven events. Image-only SAR representations behave similarly, suggesting a broader lesson: candidate events should be represented before annotation, because geographic spread is a poor proxy for useful diversity.