Selective Forgetting in Tool-Augmented Language-Model Agents
Abstract
Tool-augmented reinforcement learning can improve language-model agents when external tools are useful, but it may also change how the model behaves when tools are absent or unnecessary. We study this effect in HotpotQA and ToolACE-style tool-use training on Qwen-based models. Our results suggest a form of selective forgetting: tool-use fine-tuning does not cause uniform collapse on standard math, commonsense, or instruction-following benchmarks, but it can shift the model away from deliberative reasoning. This appears as a growing Tool-On vs. Tool-Off withdrawal gap, substantial response-length compression, reduced reflection-marker density, and transient drops on reasoning-sensitive evaluations. Qualitative GSM8K rollouts show that shorter responses can be cleaner on simple arithmetic, while quantitative forced-reasoning probes suggest that similar compression can suppress useful deliberation on harder reasoning-sensitive tasks. We argue that tool-use fine-tuning should preserve not only final-answer accuracy, but also the model’s reasoning policy.