Ungrounded: Entity Grounding Failure Drives Tool Misselection in LLM Agents
Abstract
A tool-using agent has to decide which tool to call, and whether to call anything at all. Benchmarks score whether that decision was right. We were interested in a different question: what does an agent reach for when the right tool is sitting in the catalogue but cannot be used? To make that observable we put a decoy in the registry, a tool no legitimate request should ever touch, so that misselection announces itself without anyone adjudicating individual calls. Across four studies and 13,470 trials the trigger turns out to be specific and reproducible. When the agent cannot ground an entity in the request, whether because it is unnamed (our CDN provider) or named but unfamiliar (Northbrook CDN), it treats resolving that entity as a sub-goal it must clear first, and goes looking through internal tooling to do it. fetch_url, the tool that actually serves these requests, is called in 78.1% of trials when the entity is familiar and named, 15.1% when it is named but unfamiliar, and 5.0% when it is unnamed, with the same gradient in all six models we tested across two vendors. A configuration-export decoy absorbs the traffic, firing on as much as 39% of unresolved-referent requests on the worst-affected model. Under prompt-clustered inference the effect holds in five of the six, with matched controls at zero throughout. The sixth, claude-opus-5, fired once in 720 ungroundable trials without abstaining any more often than the others, which suggests this is something training can fix.