Beyond Post-Hoc Saliency: Assessing Mechanistic Representation Interventions in a Deep Learning Weather Model
Philine L Bommer ⋅ Marlene Kretschmer ⋅ Emma Kasteleyn ⋅ Anna Hedström ⋅ Ana Lucic ⋅ Fanny Lehmann ⋅ Marina Höhne
Abstract
Deep learning weather models (DLWMs) achieve competitive global skill, but their opaque representations limit verification of whether forecasts rely on physical dynamics or statistical artifacts. While linear probes and representation interventions have enabled internal control in language models, the application to DLWM remains largely unexplored. Moving beyond post-hoc attribution-based explainability methods, we present a first systematic assessment of linear probes and mechanistic representation interventions for DLWM. We investigate the efficacy of four mechanistic interventions applied to the Aurora model across two different target cases: tropical cyclones (TCs), investigating bias correction, and El-Ni\~no-Southern-Oscillation (ENSO), investigating shifts in ENSO strength, useful for scientific validation and discovery. For TC peak wind (a localized, high-frequency event) three techniques fail despite strong probe fits, only a probe-based low-rank adaptation (Probe-LoRA) improves mean wind speed error by up to $28\%$. Conversely, for the ENSO (with Ni\~no~3.4 index), a low-frequency, large-scale event, vanilla additive steering along the probed direction cleanly intensifies and reverses El Ni\~no and La Ni\~na phases on $100\%$ of events. We release our framework in the open-source Mechanistic Interpretability and Steering Toolkit for Intelligent Climate (MISTIC).
Chat is not available.
Successful Page Load