Wording Diversity Reveals Tool-Use Robustness Failures
Abstract
Tool-using language models may act correctly under one phrasing of a request and fail under another, even when the world and intended action are unchanged. We test three models on 72 calendar and file scenarios, expressing each request through a fixed set of four manually designed phrasings and pairing states where its condition is true or false. For the other construction, testing all four phrasings reveals unsupported actions in 91.7% of Qwen worlds and 97.9\% of GPT-OSS worlds; averaged over the possible single-form choices, the corresponding rates are only 37.5% and 51.0%. For additive requests, all models take unsupported actions under every tested phrasing. Runtime checks reduce unsupported execution, but mainly by suppressing correct behavior: constrained checks permit the correct supported action in only 14.2% of Qwen cases and 0.3% of Mistral cases. Robustness evaluation should therefore test wording variation while also measuring whether useful actions are preserved.