Ranked Right, Sized Wrong: Ensemble Spread Fails Silently at Streamflow Extremes
Justin Gallant ⋅ Tyler Wilson
Abstract
Deep learning models, particularly Long Short-Term Memory networks, are increasingly used for operational streamflow forecasting, with ensemble spread providing a signal of predictive uncertainty. As climate change increases the frequency of hydroclimatic extremes, these models will increasingly operate under flow conditions rare in their historical records. We therefore evaluate whether ensemble spread remains informative under increasing flows, examining both its magnitude relative to error and its ability to rank errors. Prediction error is approximately $2\times$ the corrected spread at ordinary flows and up to $4\times$ at the highest flows, with the discrepancy increasing with flow (log-log slope $1.30$, 95\% CI: $[1.20,1.41]$). Under-dispersion is strongest in observed extreme flows beyond the training record, reaching an $18\times$ discrepancy. This is pervasive and not explainable by conditional mean bias. Area Under the Sparsification Error, a standard distribution-free metric of uncertainty ranking, indicating near-oracle skill when comparing versus error ($0.071$), yet ordering by predicted discharge alone nearly matches it. Informative uncertainty ranking can coexist with severe under-dispersion. In deployment, error ranking cannot be verified, yet the spread is available, and it is relied on as the uncertainty signal under increasingly extreme conditions, while systematically understating risk, most severely for the extreme events that matter most.
Chat is not available.
Successful Page Load