Transcoder Adapters for Reasoning-Model Diffing
Nathan Hu ⋅ Jake Ward ⋅ Thomas Icard ⋅ Chris Potts
Abstract
While reasoning models are increasingly ubiquitous, the effects of reasoning training on a model's internal mechanisms remain poorly understood. We introduce transcoder adapters, a technique for learning an interpretable approximation of the *difference* in MLP computation before and after fine-tuning. We train transcoder adapters on two pairs of base and reasoning models: Qwen2.5-Math-7B / DeepSeek-R1-Distill-Qwen-7B and Qwen2.5-32B / QwQ-32B. We find that modeling the difference in MLP computation is far easier than modeling the full MLP; adapters achieve faithful reconstruction with an order of magnitude fewer active features than typical transcoders. When evaluated on reasoning benchmarks, adapters exhibit the reasoning model's characteristic long responses and recover a large fraction of its benchmark performance. Adapter features are interpretable, achieving higher automated interpretability scores than MLP neurons. To demonstrate the utility of transcoder adapters for interpreting fine-tuning differences, we present two case studies on the 7B model pair. First, we examine the overall composition of adapter features to gain a broad overview of fine-tuning differences. Despite being specific to fine-tuning by construction, many features have activating examples unrelated to reasoning structure. Interventions confirms these features disproportionately affect benchmark performance rather than response length. Second, we study why the model says 'wait' by constructing attribution graphs whose edges flow through base model parameters. We trace hesitation to only $\sim$2.4\% of adapter features (5.6k total), finding that this behavior depends largely on base model computation. These features are necessary and sufficient for producing hesitation tokens; removing them reduces response length, often without affecting accuracy. Anonymous code is available at \url{https://anonymous.4open.science/r/transcoder-adapters-6752/}.
Chat is not available.
Successful Page Load