Beyond a Single Geomodel: Evaluating Training Diversity in Conditional Diffusion Models for Geological Inpainting
Abstract
Generative models trained on the output of scientific simulators are increasingly used as fast surrogates for uncertainty quantification, but each simulator run is expensive and it is rarely clear how many independent runs the training set actually needs. We study this data-budget question for conditional diffusion models used for geological inpainting, where the unobserved subsurface ahead of a drill bit is generated conditionally on the geology already observed behind it. We generate the training data from the same stochastic geological simulator with different random seeds. We train three diffusion models on one, two, and four realizations, respectively, while keeping the model architecture, training settings, and conditional inference parameters fixed. We evaluate generalization to a held-out realization using variogram mean absolute error and percentile coverage across eight independent training runs per configuration. The four-realization model achieves the lowest mean variogram error and percentile coverage closest to the nominal levels, although the differences among training configurations are not statistically significant across the eight repeated runs. While mean variogram error decreases with increasing training diversity, percentile coverage does not improve monotonically, and substantial variability is observed across independent runs. Our findings highlight the importance of repeated training runs when evaluating the effect of training-data composition on conditional generative models for scientific simulations.