Separating Structural from Parametric Decisions in Sub-Billion Tool Use
Abstract
We ask whether part of a small model's failure at multi-step tool use comes from the form in which decisions reach it rather than from its size. In a closed-world sandbox, a 0.6B model that solves none of 40 multi-step tool tasks when it must sequence raw primitives (L₀) solves 87.5% of them when the same tasks are re-expressed so that sequence-level (structural) choices appear as typed (parametric) arguments (L₁). The model and the executor are unchanged, both languages have identical model-independent expressivity (Esyn = Esem = 1) certified by gold replay and oracle search, and the ordering holds on 30 counterfactual variants and across five task samples. Because L₁ also cuts total decision entropy by 23%, we add a control L₀′ that restricts legal transitions without bundling: it carries less entropy than L₁ yet reaches only 15.0%, no raw-primitive grammar in a six-grammar sweep over H_tot ∈ [2.50, 3.48] scores higher, and a 2×2×2 factorial leaves ordering as the only tested constraint family associated with the residual. Replacing every name with an opaque token preserves the direction on all five renaming seeds but cuts the gap to 0.210, so informative naming supplies most of the magnitude. Candidates are enumerated and selected by argmax: decision preference, not end-to-end generation. Up a three-point ladder the gap narrows, and by 4B the lower-entropy control overtakes L₁. On a public multi-turn benchmark the procedure runs unchanged but the separation does not reproduce, since that API admits only an additive recoding. Our claims are bounded to one sandbox, one model family and enumerated decoding; the identification chain is a 0.6B result.