Deterministic vs. Model-Based Graders for Tool-Using Agents: How Much Exposure a Safety Comparison Needs
Abstract
An agent that uses tools decides on a call and then executes it. In many tool-using pipelines nothing intervenes between the two. That absence matters, because a call that runs before the information it depends on exists is not a wrong answer but a wrong action, already taken. Existing work leaves it open, either executing what the agent proposes or scoring the trajectory once it is finished. We propose a deterministic check, a test of whether each call's inputs are already available before it runs. We compare it against the natural alternative of asking the model to vet its own calls, with unchecked execution as a control. The comparison spans three models and over 15000 executed tool calls. But unsafe executions are rare, about one call in 1100, and separating the two checks at that rate would take thousands of calls in each condition, more than we or comparable studies ran. The choice does not depend on that measurement. The deterministic check executed no call whose declared inputs were unavailable, and by construction cannot when the tool schema is correct. Self-judgment carries no such guarantee, and in our runs it admitted three unsafe executions. It also cost more on two of the three models, and less on the third, so cost is not what separates the two designs. The guarantee is. The same experiments give two more results. We prove that two post-gate refinements cannot increase tool cost or round count, and observe reductions across the tested configurations. We also characterize when a domain's structure lets calls be planned in advance, so no check is needed at all.