Evidence Networks with Flow Matched References for Amortized Simulation-Based Model Comparison
Kai Lehman ⋅ Carolina Cuesta Lazaro ⋅ Niall Jeffrey ⋅ Francisco Villaescusa ⋅ Benjamin Wandelt ⋅ Sven Krippendorf ⋅ Jochen Weller ⋅ Jameson Dong
Abstract
The evidence, $\mathcal{Z} = p(x | \mathcal{M})$, is a powerful tool in Bayesian frameworks: enabling model inference, model misspecification tests, and goodness-of-fit assessment as an alternative to standard discrepancy statistics, e.g. $\chi^2$. It is also the natural quantity for model comparison via the Bayes factor. While versatile, the evidence is notoriously difficult to compute, even when the likelihood is available. In both likelihood-based and simulation-based inference, estimating $\mathcal{Z}$ provides an immediate solution to high-dimensional goodness-of-fit or model misspecification tests – using $\mathcal{Z}(x)$ as the test statistic induces a minimum-volume level-$\alpha$ test. Especially in the simulation-based paradigm, we should not be restricted to using simulations for parameter estimation. We need accurate, fast, and amortized methods to estimate both evidences and Bayes factors, exploiting the full statistical power of the simulations. We present Evidence Networks with flow-matched references. Building on the Evidence Network framework of Jeffrey & Wandelt 24, the proposed method uses neural classifiers to improve the density estimation of flow-based models. We show strict improvement in estimation of $\mathcal{Z}$ over the baseline flow-based model in all applications considered and the best performance of all density estimation methods considered in our benchmarks. We furthermore introduce an architecture that learns all combinations of Bayes factors across multiple models, as well as their individual evidences jointly, streamlining their estimation when the likelihood is intractable. We demonstrate the performance and scaling behavior of these approaches on both mock simulations and real scientific data from cosmology. We apply our framework to using the evidence as a test statistic on real data, which would be computationally prohibitive with standard likelihood-based methods, but is straightforward with our proposed framework.
Chat is not available.
Successful Page Load