Representations for the Physical Sciences
Abstract
Representation learning has become a central ingredient of modern AI, yet its scientific use raises distinctive challenges that are not captured by standard vision-and-language settings. Scientific data are heterogeneous, structured, and often generated by dynamical systems; useful embeddings must therefore respect geometry, symmetries, conservation laws, and causal structure, while remaining transferable across regimes that are expensive to verify experimentally or computationally. At the same time, the sciences offer unique opportunities for learning representations, including large-scale unlabeled datasets, physically grounded simulators, and closed-loop experimental pipelines. This workshop will bring together researchers from machine learning and the physical and life sciences to focus on four tightly scoped bottlenecks in representation learning for physical systems: self-supervision under physical constraints, transfer and its limits, sampling and closed-loop data generation, and tokenization of continuous scientific modalities. Through invited talks, contributed presentations, posters, and discussion sessions, the workshop aims to clarify the methodological foundations of scientifically useful representations and to foster a shared research agenda across communities.