Scientific Concept Bottlenecks: Self-Supervised Discovery of Interpretable Scientific Concepts from Astrophysical Time Series
Abstract
Scientific time series are generated by complex physical processes whose underlying mechanisms are often only partially observable. Although recent self-supervised learning methods have demonstrated strong capabilities in learning representations from unlabeled temporal observations, their latent spaces remain largely opaque, limiting their use for scientific discovery, where understanding how observations are organized can be as important as predictive performance. We introduce the Scientific Concept Bottleneck (SCB), a self-supervised framework that distills high-dimensional temporal representations into a compact set of emergent scientific concepts, without predefined concepts, semantic annotations, or expert supervision. SCB shapes these concepts through objectives promoting stability, independence, and predictive sufficiency. We further introduce Scientific Concept Profiles to interpret the learned concepts post-hoc through independent descriptor associations, representative temporal patterns, activation behavior, and downstream utility. Experiments on Fermi-LAT gamma-ray light curves show that several discovered concepts exhibit reproducible astrophysical associations and characteristic temporal patterns, while concept-level ablations demonstrate their contribution to downstream predictive performance. These results show that self-supervised concept discovery can turn opaque temporal representations into structured, scientifically interrogable representations.