Beyond Feature-Wise Interpretability: Discovering Causal Conjunctions in LLMs
Abstract
Sparse autoencoders (SAEs) provide interpretable feature dictionaries for language-model activations, yet model behaviors are rarely explained by a single feature in isolation. We investigate whether conjunctions of SAE features, which we call patterns, can serve as causal mechanisms of LLM behavior. We introduce RegresNaps, a differentiable pattern miner that discovers conjunctions of binary features predictive of a continuous target. A binarized autoencoder learns candidate conjunctions, a mixture-density head relates their occurrence to the target, and held-out statistical tests filter and calibrate the resulting patterns. We apply RegresNaps to Gemma-3-27B-it with GemmaScope-2 SAEs across four behaviors: factual recall, distraction, instructed deception, and refusal. We find that the causal role of a pattern depends on the target behavior. Patterns associated with factual recall mostly carry answer information, so ablating them more often disrupts a correct answer than fixes a wrong one, and only a small subset has a corrective effect. By contrast, distraction patterns support more direct mechanical control: intervening on selected patterns repairs derailments with up to a 35.5% net gain. Interventions on deception patterns restore 22.3-55.1% of deceptive answers. Refusal patterns exhibit the strongest bidirectional control, with ablation bypassing up to 99% of refusals as measured by an independent safety classifier, while activation can also induce refusal on benign prompts. Feature-level audits show that these conjunctions combine semantic state, task context, and output-level features. Differentiable pattern mining therefore provides a scalable route from sparse feature dictionaries to behavior-specific mechanistic hypotheses, while causal intervention determines whether a mined pattern causes or merely accompanies a behavior.