Do Paired Activation Explainers Speak Plain English? Testing Faithfulness in Natural Language Autoencoders
Ananth K Suresh ⋅ Arya Hariharan
Abstract
Natural Language Autoencoders (NLAs) train an encoder to emit a natural-language latent $z$ summarizing an internal activation $h$, and a decoder to reconstruct $h$ from $z$. However, high reconstruction fidelity alone does not prove $z$ is a publicly readable description of $h$, as non-faithful mechanisms could yield identical fidelity. Evaluating a six-test diagnostic battery on Qwen2.5-7B-Instruct and Gemma-3-12B Instruct, we find that readability diverges from functional meaning. We decompose readability into content faithfulness and semantic invariance and find that content faithfulness holds partially and latents carry substantial document information. However, semantic invariance fails, and while form-preserving rewrites incur minimal cost, meaning-preserving paraphrases degrades reconstruction, and counterfactual edits operate at chance. High functional lift confirms $z$ carries substantial information, but an independent reader LLM recovers only one-sixth of the paired decoder's information. We conclude that NLA latents use grammatical English to convey high-fidelity model state via a code largely private to the co-trained decoder, yielding surface readability that diverges from its functional meaning to an independent reader LLM.
Chat is not available.
Successful Page Load