Localizing and Steering a Tool-Trust Circuit in Tool-Integrated Reasoning Models
Abstract
Tool-Integrated Reasoning (TIR) models frequently discard correct code-interpreter outputs that conflict with their own prior reasoning. Recent work mitigates this behavior with an external check but treats the model as a black box. We address this problem from a mechanistic perspective, treating the accept-or-override choice as a mechanism to be located and steered. Using Edge Attribution Patching with Integrated Gradients (EAP-IG) and mean ablation, we localize the decision to a sparse circuit of 14-20 late-layer attention heads, 1-3% of each model's components, across three TIR models (ToRL-7B, SimpleTIR-7B, DemyAgent-4B). The circuit is also shared across training runs of ToRL-7B and SimpleTIR-7B, which share an architecture but not a base checkpoint, converge on 9 of their 12 highest-recovery heads and overlap at Jaccard 0.70. Finally, two training-free, token-free interventions derived from the circuit, head scaling and a difference-of-means steering vector applied only after the tool result, shift the model's logit difference toward the tool on held-out natural Tool-Ignored traces, without changing the prompt or adding tokens.