Fine-Tuning a Tool-Using Agent’s Source Preference Removes Prompt-Level Control
Abstract
An agent that calls a tool answers from two sources that need not agree, the facts in the model's weights and the record the tool returns. We show that a parameter-efficient adapter drives tool-following to either extreme and removes output-level control under the tested prompts. Two QLoRA adapters trained from one base take tool-following to 0.996 and to 0.004 on held-out data, from 0.629, and transfer to relation types they never saw; across five system prompts, one of which instructs the opposite behaviour, not one held-out instance changes hands. An SFT control adapter with the same target templates but no fixed source preference helps distinguish the losses. It loses the surface-form instructions the other two lose, so that loss also occurs without teaching a fixed source preference, but the memory-first prompt still moves it on 31 instances, against none for the fixed adapters. Candidate scoring shows strong attenuation of prompt sensitivity, with a log-odds span between the tool-first and memory-first prompts of 13.73 nats for the base model and 1.16 for the tool adapter. The trade-off is a cliff. At 25 training conflicts the behaviour is set while disclosure prompting retains a large effect, and by 50 the span across policies is 0.004. All of this rests on 1284 controlled counterfactual records from PopQA, a 680-item split of which yields 3398 conflicts across five open models, a conflict defined per model against its own answer, on which tool-following ranges from 0.604 to 0.956 while the models take a correct record over their own answer on 409 of 411 pooled cases. The wider cost is real but bounded. IFEval prompt-level strict accuracy falls from 0.719 to 0.592, nine of those thirteen points already spent by the control.