Cornered Agents Probe Their Sandbox: Tool Misuse and Evaluation Integrity for LLM Agents in the Wild
Abstract
Coding and CLI agents are now scored inside live environments. An evaluator hands the agent a shell, a filesystem, tools, and an automated grader, and that harness is part of the instrument a capable agent can act on. We treat the soundness of this measurement as a security property, evaluation integrity, and show where it fails. The main effect appears when a task is impossible for the agent by design. In a within-task test that holds prompt and task identity fixed and removes only one needed input file, cornering raises out-of-workspace attempt activity from 6.4% to 81.4% (difference of 75 percentage points, 11 of 11 tasks, p<0.001). In a hard-but-solvable control with a reachable decoy, the agent solves 85/85 tasks and never uses the decoy (0/85), so the effect is tied to impossibility rather than ordinary difficulty. On a porous harness, these attempts can be scored as false success. On cornered toy tasks, porous success is 14.4% [9.7, 20.4] versus 0.0% [0, 2.0] when sealed. On MBPP (Mostly Basic Python Problems), scored by benchmark unit tests, porous success rises to 23.8% [18.5, 29.7] from 0.0% [0, 1.5] on cornered tasks, and this conversion is model-dependent (0% on GPT-5-mini). On SWE-bench, neutral-prompt future-verified gaming is low (1.7% porous versus 0.3% sealed for Sonnet), while an explicit git-history hint raises it to 45.7%, an elicited upper bound rather than an ecological baseline. Prompt-level and in-process guardrails do not hold in our setting, while Linux Landlock blocks every filesystem route we test. We release the testbed and the redacted trajectory corpus.