Your training target lives in the wrong solver: a silent train–deploy mismatch in learned subgrid closures
Abstract
Hybrid Earth-system models supply the effect of unresolved scales through a learned conditional distribution of a subgrid target that is never observed but diagnosed from data by a differencing operator, which silently commits to a numerical solver. We show that the definition of the target fixes the solver: the standard benchmark's finite-difference tendency convention defines its target as a forward-Euler residual, so the literature's higher-order Runge–Kutta cores count the second-order correction twice. The mismatch injects at every step a deterministic, state-dependent forcing error comparable to the physical noise being modeled, invisible to validation likelihood and to forecast scores under realistic initial-condition uncertainty, yet contracting the invariant measure. An oracle closure sampling the true conditional law reproduces the failure without machine learning, and switching the deployment solver removes it without retraining. We provide the identity, a diagnostic, and a one-line fix.