Ordering or Pruning? What a Tool Shortlist Buys a Small Agent
Abhinav Sood ⋅ Liya George
Abstract
Retrieving a shortlist of tools does two things at once: it puts likely tools near the front, and it removes the rest. We separate them, across 120,000 tool calls on five models from 0.5B to 7B and catalogues of 8 to 128 tools. Where the shortlist holds every tool, so retrieval can only sort, sorting is worth $+12.6 \pm 4.4$ points to a 0.5B model and nothing measurable to a 7B one. For the 0.5B that is about half of what full retrieval delivers at the same catalogue size, so much of what a shortlist buys the smallest model is position rather than brevity. Which width to choose also turns on model size: widening from the retriever's top four to its top sixteen costs a 0.5B model 11.6 points and gains a 7B model 7.3 at a 128-tool catalogue. That gain is recall, not behaviour. On the queries where both widths already hold a correct tool, no model prefers the wider list and the small ones are still hurt by it. Sorting is also the expensive edit to serve: it moves the first changed byte to the front of the prompt and forfeits the prefix cache, so a re-sorted turn costs a full cold start where leaving the block alone costs 11 to 75 times less on vLLM. The serving penalty grows with model size as the accuracy gain shrinks, so the recommendation depends on size: at 0.5B sorting buys 12.6 points for about a quarter of a second per turn, and at 7B it buys nothing for three seconds.
Chat is not available.
Successful Page Load