Small Models Learn to Call, and Stop Declining: Abstention Loss in Tool-Use RL at 1–2B
Abstract
Reinforcement learning is the usual way to make a small language model a competent tool caller. What it does to the model's willingness to not call is rarely measured. In two independent 1–2B training setups, RL improved tool calling under the executable checker while abstention fell in both pipelines: BFCL irrelevance detection dropped by 26 and 32 points. The optimized metric moved the other way, rising 13.6 points under the checker and staying within noise under the judge. One setup was scored by a 122B LLM judge, the other by a deterministic executable state checker, so the loss also occurs in a pipeline using an executable verifier, though the two setups are not a controlled judge-versus-checker comparison. Neither setup's training data contained an example whose correct answer was to call nothing. We then tried to fix it, and ran the ablations that say which part works. Adding decline turns to the data and a light penalty for calling on them kept abstention near its starting value through step 140 at no measurable cost on the target metric; a heavier penalty only postponed the loss. A positive credit for a correct no-call, which we expected to be the part that mattered, was not: a run identical except for the credit was at or above it at 14 of 18 checkpoints. Removing the penalty instead, with the decline turns still in the data, collapses abstention (below 23% from step 60 on, ending at 13.8%), worse than never adding the decline turns at all, and it trails the control on multi-turn too. In these runs the price on the erroneous call is the operative term: coverage alone is not enough, and the positive credit adds nothing measurable at the weight we tested. We recommend evaluating small tool-use models on abstention next to task success, and putting a cost for calling when the right move is to decline into the training signal.