Cross-simulator transfer with self-supervised representations in 21cm cosmology
Abstract
Simulation-based inference is limited by model misspecification: neural summaries trained on one simulator degrade on data from another, or from a real instrument. We study this in 21cm cosmology, where the brightness-temperature field from the Epoch of Reionization is targeted by the Square Kilometre Array (SKA). We pretrain a Vision Transformer once with a self-supervised Joint Embedding Predictive Architecture (JEPA) objective, label-free, on cheap semi-numerical simulations, freeze it, and transfer to a physically distinct hydrodynamical radiative-transfer simulator with a disjoint parameter set, under realistic instrumental noise. The frozen encoder, shown neither the target simulator, its parameters, nor any noise, matches or exceeds a supervised network retrained directly on the noisy target, and retains calibrated posteriors. A controlled ablation holding architecture and pretraining data fixed shows that the self-supervised objective, not data scale, is what enables this robustness.