What Makes Diffusion-Based Weather Downscaling Work? A Controlled Comparison of CorrDiff- and Patch-DM-Style Models
Kai Kwan LAU ⋅ Gary P. T. Choi
Abstract
Diffusion models are becoming the default for generative weather downscaling, turning coarse climate fields into the kilometer-scale ensembles behind heat- and wind-risk planning, but published systems entangle the architecture, the diffusion formulation (DDPM vs. EDM), and the evaluation protocol. We disentangle them with a controlled factorial, with shared backbone, pipeline, and metric suite, on a perfect-model ERA5 benchmark: CorrDiff (two-stage, residual) and Patch-DM (single-stage, patched; to our knowledge the first atmospheric use of its feature-collage scheme), each trained under both formulations. The formulation acts through an interaction with the diffusion target. On residual targets at $\times 8 \to 128^2$, DDPM matches or beats EDM on probabilistic skill, calibration, and member spectra at half the sampling cost. On full targets, DDPM is seed-unstable at $128^2$ and collapses at $256^2$ (5.5 K error, worse than bicubic) despite converged losses, while the same architecture under EDM is the best configuration overall. Replacing only the prediction target with $v$ removes the collapse at unchanged cost ($5.9 \to 0.60$ K); equal-NFE controls change nothing: the failure is the $\varepsilon$-parameterization at high noise, not "DDPM" or the sampling budget.
Chat is not available.
Successful Page Load