Was the Feature About the Concept? Four Faithfulness Tests for Sparse-Autoencoder Attributions, and a Case Where All Four Fail
Abstract
When foundation models are used for scientific discovery, internal features may be interpreted as evidence that the model has learned a meaningful concept. But a feature can activate strongly on examples containing a concept without being specific to that concept. We propose four inexpensive faithfulness checks for sparse-autoencoder (SAE) attributions: compare against matched controls, compare interventions against random features, test whether the features are specific to the relevant domain, and repeat the attribution at a different read-out site. We apply these checks to a standard top-k-by-activation attribution for insecure cryptographic code in Gemma-2-2B-IT, a proxy domain chosen because it supplies the matched controls the tests need and a human-labeled outcome. All four fail. None of the 84 attributed features responds more than 1.5× as strongly to insecure as to matched secure prompts; the same features respond nearly equally to ordinary Python; rank-matched random features do as well downstream in every pre-registered comparison; and re-selecting the features at the pre-suffix read-out site replaces 10–14 of the 16 features. We observe the same pattern with a 4× wider dictionary, on SQL injection, at the set level on a second architecture (Llama-3.1-8B with Llama Scope dictionaries), and for an El Niño versus La Niña concept posed to the same model in text, where the content-bearing sites yielded only weakly discriminative features. In the cryptographic case, more discriminative features exist at other read-out sites, but intervening on them produces only a small, inconsistent effect. High activation alone is weak evidence that an SAE feature represents the concept assigned to it; the four checks are simple controls to run before an attribution is treated as scientific knowledge. The tests are the contribution; the case is the warning.