Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry
Abstract
The rapid development of foundation models (FMs) finetuned for atmospheric chemistry prediction offers a novel target for mechanistic interpretability work. We study the mechanisms of an FM fine-tuned for atmospheric chemistry, specifically Microsoft's Aurora model. Internally, Aurora's representations remain largely organized around meteorology encountered during pretraining, with little chemistry-specific structure. We train AuroraScope, a suite of 48 sparse autoencoders, on the residual stream of the model's transformer operator to identify internal features which causally control the chemical forecast but that do not map cleanly onto individual atmospheric processes. Our work provides a framework for understanding how AI forecasting systems learn atmospheric chemistry from reanalysis data, and as these models are increasingly positioned to inform environmental policy decisions, we argue that composition forecasts should also be judged by their internal mechanisms rather than by benchmark skill alone.