Evaluating Request Assumptions in Tool-Using Language Models
Abstract
A tool call can be legal while relying on a false assumption conveyed by the user's request. We evaluate Qwen3-8B, Mistral-7B, and GPT-OSS-120B on 72 synthetic calendar and file scenarios, holding state and requested action fixed across backgrounded and asserted formulations. All three models execute every unsupported additive request across four backgrounded forms. Explicit assertions reduce these actions, but also reduce correct supported execution to 33.3\% for Qwen and 75.0\% for Mistral. Responses to other vary sharply across forms and models; GPT-OSS's unsupported action rate ranges from 0\% to 93.8\%. On the same benchmark, constrained generation improves Mistral's runtime-check validity from 45.5\% to 99.0\%, yet retains only 0.3\% correct supported actions. These results show wording-sensitive action under a stipulated interpretation, rather than a uniform effect of backgrounding. They also distinguish producing well-formed checks from extracting conditions that support useful execution.