Same Label, Different Feature: Explainer Dependence and the Interventional Limits of SAE Explanations
Abstract
Sparse-autoencoder (SAE) features are often selected as steering, monitoring, or unlearning handles by reading their automatically generated explanations. Yet these explanations are usually evaluated against activations, leaving unclear whether identically labelled features have similar causal effects. We introduce a dyadic intervention audit that tests this assumption by intervening on feature pairs assigned the same exact explanation and comparing their downstream-effect directions with same-pool random pairs on identical contexts. Across four public explainers applied to the same 16,384 Gemma-2-2B layer-20 features, exact collision rates (one explanation assigned to multiple features) range from 3.1% to 72.2%, and only 0.05–1.05% of one substantive explainer’s colliding pairs also collide under another. A post-hoc audit finds that label form accounts for a measurable part of this explainer dependence without establishing its cause. At layer 12, the raw same-label association is +0.265 matched-random SD (95% CI [+0.129, +0.401]), but prespecified adjustment for decoder geometry and activation rates reduces it to +0.075 SD ([−0.063, +0.205]). A frozen layer-20 measurement is also weaker at +0.057 SD ([−0.061, +0.175]), under a corpus draw that was not held fixed. Within this evaluated setting, shared labels prioritize candidates. They do not establish causal interchangeability. More broadly, automated feature explanations should be treated as explainer-specific hypotheses that require direct intervention tests before model control. Code: https://anonymous.4open.science/r/sae-collision-7ECE