Interpretable Units in Chemical Language Models: A Two-Model Audit of Neuron Interpretability and Causal Control
Shreyash Goli ⋅ Ishan Gonehal
Abstract
Chemical language models predict molecular properties from SMILES strings, but a persistent skeptical reading holds that they exploit surface string statistics rather than chemistry. Interpretability work on these models has been almost entirely correlational and single-model, and correlational evidence cannot separate the two hypotheses: SMILES token count correlates with molecular weight at $R^2 = 0.92$, so a neuron that counts tokens passes a descriptor-alignment test as convincingly as one that detects a functional group. We audit two architectures that differ by $4\times$ in depth, ChemBERTa (3 layers, 1,392 MLP neurons) and MolFormer (12 layers, 9,216), against matched sparse autoencoders trained on the same corpus. Residualising every descriptor against token count removes 22–39% of the units that pass a specificity gate, in both models, and leaves no ChemBERTa neuron aligned to molecular weight. The survivors are genuine substructure detectors: 3,226 of 3,248 motif-aligned units beat an exact permutation null, exceeding a prevalence-matched chance floor by $0.21$–$0.37$ F1. Single-neuron control is null, and a naive faithfulness metric inverts it, turning an apparent $+0.19$ lift into $+0.05 / -0.16 / -0.05$ once scored against each neuron's true property-correlation sign; this analysis needs parsed label directions and so is ChemBERTa-only. Coordinated interventions steer, but at a fixed intervention budget of $K = 200$ units the raw-neuron basis does not survive scale. On a probe-free readout, ChemBERTa's raw neurons move held-out molecules three to six times further than their own sign-shuffled control; MolFormer's are the size of that control on two tasks and, on the third, larger but pointing the wrong way. Matched SAE latents beat their control in both models. Counting significance at $p < 0.05$ across three datasets and four readouts tells the same story, at 9 of 12 and 1 of 12 for raw neurons against 12 of 12 for latents in both. Both models encode real chemistry; between a fifth and two fifths of their apparent interpretability is string length; and at a fixed intervention budget of $K = 200$ units, the raw-neuron basis loses causal purchase with depth while a matched sparse basis does not.
Chat is not available.
Successful Page Load