Verifying Tool-Using Agent Interventions: When Do Training and Runtime Controls Help?
Siddharth Ramakrishnan
Abstract
Frontier models can call tools very well, but using them for every routine operation can make an agent unnecessarily expensive. We study whether bounded, high-frequency tool calls can instead be assigned to a smaller model like Qwen2.5-1.5B and compare two paths: supporting the unadapted model with an inference-time harness or specializing it through supervised fine-tuning (SFT). On a held-out 100 example BFCL subset, the unadapted model reaches 18.0% exact match on its own and 36.0% with our full harness. SFT is substantially stronger: verified training demonstrations yield $74.7 \pm 0.6%$ across three seeds. Demonstrations generated by GPT-5.6 Sol from training requests and tool schemas alone yield $70.3 \pm 2.3%$. These results show that standard SFT can create a useful small tool-calling specialist, while runtime scaffolding offers a weaker fallback when post-training is unavailable. They support multi-model systems that route bounded calls to a small specialist and reserve frontier models for harder cases.
Chat is not available.
Successful Page Load