AegisBench: Auditing Aerial Search-and-Rescue Person Detection Under the Conditions It Is Deployed In
Abstract
Aerial person detection is being adopted by public emergency-management agencies to find missing people faster than ground teams can search. The benchmarks that justify that adoption were collected in calm, clear conditions, while the missions that trigger deployment happen during floods, wildfires, storms, and at night. We treat this as an evidence-standards problem and build an audit for it. AegisBench evaluates aerial search-and-rescue person detection under nine physically motivated visual corruptions spanning four disaster families, each at three severities calibrated against a declared, machine-verified image statistic. We evaluate three architecturally distinct detectors across 168 conditions on two datasets covering the high- and low-altitude capture regimes, scoring every model with one shared evaluator at operating points frozen on clean validation data, as a deployed system would be. Robustness is highly uneven, and uneven for the same corruption across datasets: heavy rain costs about 3.6 points of recall at the highest severity on SARD but is among the most damaging corruptions on HERIDAL. Low light is worse and universal. Recall reaches exactly zero for all three architectures on both datasets, with bootstrap confidence intervals degenerate at [0.000, 0.000]; a threshold-independent mAP analysis confirms the detections are absent rather than below threshold, and a check on real night drone imagery reproduces the collapse. A localization-stability analysis shows surviving detections stay accurately placed, so the failure is a detector going blind rather than misplacing what it finds. The failure is silent: nothing in the system's output distinguishes an empty frame from a frame whose occupants it cannot see. We release the benchmark, corruption engine, and evaluation pipeline as an audit instrument for agencies procuring these systems.