OroPrecipBench: a km-scale benchmark for spatial precipitation downscaling over complex terrain
Abstract
Precipitation downscaling over complex terrain is essential for resolving water availability, flood risk, and extremes, but current evaluations rarely test whether models learn orographic structure rather than only pixel-wise accuracy. We introduce OroPrecipBench, a km-scale benchmark for orographic precipitation downscaling across five U.S.\ terrain regimes. The benchmark provides a shared-domain dual-track protocol: a perfect-model track that coarsens 3\,km HRRR analyses to isolate controlled super-resolution skill, and an observation-grounded track that downscales 25\,km ERA5 to 4\,km PRISM to test practical robustness under source-target shift. We also introduce PEP-D, a terrain-conditioned diagnostic that measures errors in elevation--precipitation profiles, complementing pointwise, spatial, and extreme-precipitation metrics. Evaluating ten statistical, deterministic deep-learning, and generative baselines shows that strong controlled super-resolution does not reliably translate to real-case or cross-regime downscaling. Models that reduce pixel-wise error can still distort terrain-dependent precipitation and upper-tail rainfall, and terrain-conditioned structure is especially fragile under transfer. OroPrecipBench exposes these failure modes and motivates source-shift-robust, terrain-constrained, and tail-aware downscaling methods. Benchmark is available \href{https://anonymous.4open.science/r/OroPrecipBench-612D}{here}.