How much conditional calibration does marginal conformal prediction buy an AI weather ensemble?
Abstract
A 90\% prediction interval should contain the outcome about nine times out of ten. Intervals from AI weather ensembles fall short of this, and a recent online conformal method restores the average coverage with a distribution-free guarantee. Decisions based on weather intervals are made mainly on days when the forecast indicates a possible extreme, and a guarantee about the average does not constrain coverage on those days. We measure what the average-level correction leaves behind for a 20-member ArchesWeatherGen ensemble of 2\,m temperature at 5-day lead, over a full year of daily forecasts, by grouping days according to the fraction of members that exceed a local climatological threshold. The raw intervals cover 69.7\% of outcomes in the highest-warning group and 73.8\% in the lowest. The average-level correction removes most of this difference. Coverage in the highest group reaches 89.6\% globally and 88.1\% over land, and the correction widens intervals most in that group. Independent per-group corrections would address the remainder, but the measured adaptation speed of the method implies about seventeen years of daily forecasts to calibrate the rarest group, and coupled group-conditional methods remain to be evaluated in this setting. We also find that adapting the quantile level on the ensemble's own distribution, instead of padding additively, does not reach the target at this ensemble size. The interval saturates at the ensemble range on 95.6\% of forecasts and coverage settles at 84.4\%.