Beyond Tool-Call Count: Cost-Aware Tool-Use Routing for Black-Box LLM Agents
Abstract
Large language model agents often invoke external tools even when they can answer directly, incurring unnecessary cost. Existing tool-routing methods commonly measure efficiency through tool-call count, but this is an incomplete proxy for realised execution cost: costs vary substantially across tools and tasks, and tool use can occasionally be cheaper than answering directly. Moreover, many lightweight tool-need predictors rely on hidden states unavailable from black-box API models. We introduce a cost-aware extension of Probe & Prefill that combines predicted tool necessity with an estimate of the marginal cost of tool use. Off-the-shelf embeddings replace target-model hidden states, enabling black-box deployment, while a user-controlled weight trades off tool need and cost. On When2Tool, cost-aware routing improves the accuracy--cost Pareto frontier over Probe & Prefill by 14.7\% in hypervolume, with substantially smaller gains on the accuracy--tool-call frontier. A simple environment-level cost average performs comparably to learned task-specific regression and counterfactual estimate. These results show that call-count-based evaluation can obscure important differences between routing policies, while simple calibration of realised cost provides a richer signal for both evaluating and controlling tool use.