Where Can We Trust Neural Emulators? Simulator-Relative Validity Profiles
Abstract
Scientific workflows increasingly replace expensive numerical simulators with neural emulators. Since simulators are imperfect representations of reality, two distinct discrepancies arise: simulator-to-reality and emulator-to-simulator. Low emulator error therefore cannot establish physical validity, but measuring emulator-to-simulator discrepancy is necessary for assessing pipeline reliability. Focusing on this layer for initial-boundary value problems (IBVPs), we introduce a simulator-relative probing protocol. By sweeping factors such as physical coefficients, initial-condition parameters, or numerical fidelity, we map emulator error along controlled simulator axes and obtain thresholded validity domains. We use this diagnostic to study out-of-distribution generalization, continual learning, and active simulation acquisition. Across linear and nonlinear PDEs, we demonstrate that the validity set does not simply coincide with the convex hull of the observed configurations. Instead, validity forms localized, directionally sensitive basins around observed regimes. Consequently, interpolation failures persist even within the geometric span of the training data, and extrapolation sharply degrades. Our results motivate evaluating neural emulators through simulator-relative operating profiles rather than aggregate test errors alone.