Model-Free Assessment of Simulator Fidelity via Quantile Curves
Abstract
As generative AI models are increasingly used to simulate real-world systems, quantifying their “sim-to-real” gap is critical. We study this gap across a population of input settings, called scenarios, e.g. survey questions or operating conditions. The population-level quantities defining the real-world and simulated system are only observed through finite, often heterogeneous, samples. Consequently, the population-level discrepancy cannot be directly computed, and standard predictive inference methods that target observable outputs are ill-suited. We propose a model-agnostic framework that constructs confidence sets for the latent parameters, forms a conservative proxy for the sim-to-real discrepancy, and estimates its quantile function across scenarios. The resulting calibrated risk profile allows for inference on a new scenario, tail-risk summaries such as Conditional Value-at-Risk (CVaR), and principled comparisons across simulators. Our method applies to a range of output spaces, including categorical survey responses and continuous outcomes. We demonstrate its utility by evaluating the alignment of four major LLMs with human populations on the WorldValueBench dataset.