"All Models Are Right": The Blind Spots of Posterior Predictive Checks in Simulation-Based Inference
Abstract
"All models are wrong, but some are useful" [Box, 1979]: simulators play a central role in explaining observed phenomena from the real world, and the question is not whether they are wrong, but whether they are wrong enough to matter. Standard practice is to use simulation-based inference (SBI) to fit a neural posterior estimator (NPE) and assess the model fit via posterior predictive checks (PPC) -- simulate replicas from the inferred simulator parameters and compare them to the observed data. Two things can go wrong: the posterior approximation itself might be unreliable, or the diagnostic used for the comparison could be blind to exactly the mismatch it's supposed to catch. We investigate and formalize the second: for which summaries does a PPC always pass? The intuition is that any summary informative only for the simulator creates a circularity -- it maps data to its closest representation in the simulation space and the posterior for that representation reproduces it exactly. First, we prove this in closed-form for the case of exponential families and conjecture it holds for the NPE's own summary network too -- which we verify experimentally on the pendulum task from SBI literature. So be aware: using the NPE embedding for PPCs can make it look like "all models are right". Then we show that escaping circularity requires domain expertise and/or combining multiple complementary summaries and illustrate this by comparing increasingly more complex cosmology simulations of weak lensing maps.