Auditing Tool-Call Evaluation for Local Small Language Models
Abstract
A tool-agent score can hide three outcomes: whether execution reaches a submit state, whether the submitted state is correct, and whether the executed trace is exact. For each model–task input, we generate the initial response once. We then either reject multiple calls, run only the first call, or run all calls; each choice also fixes the subsequent history and schedule. The graph-addition grid contains 144 initial responses and 432 resulting runs across three local 8–10-billion-parameter model tags. Ministral reaches submit on 0/48 reject-multi runs and 48/48 runs under either adapter, but execute-first is exact on 48/48 versus 39/48 for execute-all. Qwen has no within-input difference; Llama never submits. Two adaptive controls and one ceiling diagnosis show different task- and model-specific patterns. Thus an evaluator bundle can change coverage and correctness differently. We support only finite, configuration-specific comparisons, not a universal policy ranking or a topology-causal claim.