Self-Supervised Pretraining for Instrument-Robust Representations of Scattering Data
Abstract
Machine learning models for scattering data are typically trained on simulation because labeled measurements are scarce. However, instrument effects absent from the simulation cause these models to transfer poorly to measurements. Existing approaches learn task-specific robustness to these effects during supervised training that must be relearned for every facility and labeled dataset. Here, we pretrain an encoder self-supervised on a large dataset of pairs of spectra that share a structure and differ in instrument parameters. Pretraining this way yields a representation in which structural targets remain linearly readable when the training spectra come from a different instrument or material class than the test set, and when labels are few. In the regime that experimental pair distribution function analysis occupies, with few labels and heterogeneous instruments, a frozen encoder and a closed-form probe outperform supervised training. Fine-tuning, where labels permit it, adds accuracy at some cost in robustness. One pretrained encoder serves every task, pointing toward the possibility to train a foundation model for scattering data.