Ungrounded Enforcement: When Agent Benchmarks Grade on Conditions Their Specifications Do Not Determine
Minhao Li ⋅ Zijian Liu ⋅ Huizhu Lyu ⋅ Hongwei Cheng ⋅ Huadong Guan
Abstract
Agent benchmarks communicate tasks in natural language but grade them with hidden programmatic oracles. We ask whether the public task specification determines what the oracle will accept. We call an enforced condition ungrounded when it is neither stated in nor entailed by the materials available to the agent, and introduce Specification Recoverability $\mathrm{R}(c)$ as an empirical audit measure. We conduct case studies of one task from each of $\tau$-bench, Terminal-Bench 2, and SpreadsheetBench, using ten independent readers per task and a blinded entailment matcher. The clearest failure is a direct specification–oracle conflict: $\tau$-bench permits a refund to either the original payment method or a gift card, whereas the oracle accepts only the former. We also identify hidden representation and execution constraints, including an unannounced fingerprint format and a spreadsheet evaluator that can reject the requested live formula. All three audited oracles additionally use an all-or-nothing aggregation rule that none of thirty readers recovered. The audit motivates a simple remedy: publish the conditions and aggregation rule that determine a task's score while keeping test instances hidden.
Chat is not available.
Successful Page Load