RTCBench: Evaluating Tool Use of Large Language Models Beyond Oracle Access
Abstract
Large language models (LLMs) have demonstrated promising capabilities in tool calling. However, existing benchmarks predominantly evaluate models in an idealized setting where ground-truth tool definitions are provided in context. This assumption diverges significantly from real-world deployment scenarios, where the model must first identify user intent and retrieve relevant tools from a large, heterogeneous tool repository before invoking them. In this work, we introduce a Retrieval-augmented Tool-calling Benchmark (RTCBench). In RTCBench, we collect 1,840 tasks from existing benchmarks and augment them with two tool environments: 1) None: no available tools provided in the prompt and 2) Distractor: strong distractor tools, while the target tools are excluded. By removing oracle tool access, RTCBench compels LLMs to actively retrieve relevant tools from a large-scale repository via a unified search tool before making tool calls, thereby faithfully simulating the end-to-end complexity of real-world tool-use pipelines. Through comprehensive evaluation of diverse LLMs ranging from 4B-parameter models to proprietary frontier models, we identify two core deficiencies: unreliable tool discovery, where models fail to initiate effective retrieval actions, and post-retrieval grounding failures, where models cannot reliably select and invoke the correct tools from semantically similar retrieved candidates. Furthermore, to mitigate these issues, we release 13,278 retrieval-augmented tool-calling expert trajectories with high-quality chain-of-thought reasoning for LLM post-training. Experiments show that a 4B-parameter model post-trained with GRPO improves its tool-calling accuracy on RTCBench from approximately 47.9\% to over 85.8\%, approaching the performance of the state-of-the-art model Gemini-3.1-Pro.