What Does a Sparse Autoencoder Feature Do? A Weight-Based Account
Abstract
Sparse autoencoders (SAEs) have emerged as a powerful technique for decomposing language model representations into interpretable features. Current interpretation pipelines infer feature semantics from activation patterns, implicitly assuming every feature has a semantic explanation while overlooking the computational roles features inherit from the SAE training objective. We introduce a weight-based interpretation framework that requires no activation data and grounds claims causally through targeted feature-ablation interventions. Across Gemma-2 and Llama-3.1, three depth-dependent signatures show that SAE features inherit the base model's geometry: semantic features form a U-shape under tied embeddings but bifurcate when untied, attention participation peaks mid-layer, and the two roles couple oppositely for input- vs. output-oriented features. This weight-based view supplies the causal half missing from activation-based interpretability.