Exploring Tool-Use Awareness with LLM Agents
Abstract
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks reward execution success while largely overlooking tool-use awareness: the ability to discern whether a problem requires external tools or is solvable from internal capability. The distinction is operationally consequential, as unnecessary or omitted tool calls increase latency, cost, and exposure to cascading errors. To address this, we introduce KAPRO (Knowing–Acting Quadrant PRObe), an evaluation framework that decouples explicit judgments of tool necessity (Knowing) from autonomous tool-use behavior (Acting). We further construct KAware, a dataset rigorously partitioning tasks into external, internal, and hybrid settings to probe agents' epistemic boundaries. Across 18 LLM agents, we find that tool-use failures follow distinct setting-dependent patterns. Agents mainly underuse tools when external functionality is required, frequently make errors on both probes for hybrid tasks, and substantially overuse tools when tasks are internally solvable.