What Survives Verification: Steering an Occupation-Gender Stereotype Feature in Two Small Language Models
Shivam Verma
Abstract
A growing family of bias-mitigation methods works by naming a specific internal object: ablate this direction, clamp that sparse-autoencoder (SAE) feature. We ask how much of such a recipe survives verification in one model, and transfer to a second, at a scale reproducible on a laptop. Using public SAEs on Llama-3.2-1B-Instruct and LFM2.5-230M we locate occupation→pronoun stereotype features by requiring two independent lines of evidence per candidate: activation selectivity across occupation-matched templates, and a decoder direction that pushes the target pronoun. Subtracting one feature (f32258) in Llama moves the he–she log-odds on male-stereotyped occupations from +2.76 [+2.35, +3.15] to +0.09 [−0.07, +0.25], an interval containing parity, and the effect holds on 24 occupations and two template families never used for discovery. Projecting the direction out removes almost all of the gap (+2.52 → −0.20 on held-out occupations), so the behaviour depends on it and does not merely correlate with it. Our controls then revised two conclusions we would otherwise have published. At matched intervention norm the feature beats difference-in-means and random directions, but the raw unembedding $W_U$[" he"] − $W_U$[" she"] beats the feature, reaching parity at +1.4% WikiText loss against +2.3%. What the SAE buys here is a direction we can name, not a better one. And a discovery ranking is not a causal ranking: in LFM2.5-230M the top-ranked masculine candidate is inert while the rank-2 candidate swings the bias by 3.16 log-odds, which overturns a cross-model claim we had drawn from rank 1 alone. We also report two silent SAE failure modes that invalidated our first results, corrupt normalisation statistics and beginning-of-sequence contamination that inflates reconstruction error 369×, together with the two-number check that catches both.
Chat is not available.
Successful Page Load