Leveraging Imperfect Simulators for Epidemic Forecasting: A Case for Standardised Benchmarking
Abstract
Synthetic data from simulators such as agent-based, compartmental or metapopulation models can be used to overcome limitations of real-world infectious disease data, which are often noisy, incomplete, and limited in the range of epidemic scenarios they capture. However, existing studies show that models which perform well in-distribution can fail to generalise to outputs from unseen simulators or to real-world data. To ensure claims of forecasting skill reflect genuine performance rather than skill specific to a simulator or metric, we propose a minimum reporting standard for evaluating forecasts on simulated data. As forecasts increasingly inform public health policy, such standards are essential to prevent overconfident models from misleading decision makers during outbreaks.