The Parser Is Part of the Agent: Auditing Silent Tool-Call Repair in Small Language Models
Pranav Kompally ⋅ Sibi Chakkaravarthy Sethuraman
Abstract
Small language models (SLMs) often emit tool calls that are semantically recognizable but violate the declared execution format; harnesses may reject them, retry, or silently repair and execute them, changing both the future trajectory and the measured agent. We isolate the execution policy on paired deterministic, stateful tasks holding model, prompt, decoding seed, tool budget, and initial state fixed. For Qwen2.5-1.5B, permissive repair inflates exact task success by 2.5-25 points across all eight resampled seeds (mean 15) and by 47.5 points in the greedy evidence run (7/40 to 26/40; 95% task-bootstrap interval $[0.325,0.625]$, exact McNemar $p=3.8\times10^{-6}$), while apparent model calls fall from 397 to 246; the greedy magnitude sits above the entire sampled range, inflated by perseveration on one malformed string. Repair is not purely additive: the greedy run loses no task, but every sampled seed loses 1-4. Qwen3-1.7B is an exact null control: both policies solve 27/40 with no repair. Repair therefore does not generically improve an SLM; it can manufacture a large, model-specific gain. Reclassifying every call across all ten strict-v2 runs collected here (seven families) reveals zero schema violations among 1,572 tool-attributed argument-bearing attempts, while four models across three families violate specifically on the two empty-argument actions (finish, list_folders). Two further models fail at the turn protocol instead, batching whole plans into single turns; recovering the calls they never dispatch reproduces the same asymmetry (1.4% vs. 32.6% violation rates). A preregistered ablation shows the boundary is causal, not an artifact of our schema: giving finish one required string argument drops its violation rate from 93% to 0% while the still-empty-argument list_folders does not improve, and strict success rises 7/40 to 29/40. This one-line schema change recovers more capability than permissive repair, with no malformed call ever executed.
Chat is not available.
Successful Page Load