Task-Agnostic Self-Supervised Learning for Physical Dynamics
Abstract
Physics foundation models are commonly pretrained by forecasting future states, tying representation learning to a specific downstream task. We investigate whether self-supervised learning from partial observations can instead produce task-agnostic representations that capture the hidden dynamics of physical systems. Using a shared ViT backbone for heterogeneous 2D and 3D systems, we study masked reconstruction in physical space and JEPA-style prediction in latent space. To assess what these representations encode about the underlying dynamics, we freeze the encoder and measure how well governing physical parameters can be recovered from its features. To explicitly target properties of the dynamical system that persist across an entire trajectory, we introduce \emph{regime fingerprinting}, which exploits the fact that an early partial view and a broader view of the same trajectory are generated under the same physical regime, and trains their representations to remain identifiable across time. Combined with latent-space prediction, regime fingerprinting yields the strongest recovery of governing parameters in our experiments and compares favorably with forecasting-pretrained baselines. These results suggest that self-supervised learning can capture trajectory-level physical dynamics without using forecasting as a pretraining objective.