DetourBench: Trajectory-Level Evaluation of Interactive Agents with Exact Minimum-Action Bounds
Hanlin Tian ⋅ Sihan Zhu ⋅ Zhao Yang ⋅ Yuxiang Wang ⋅ HU ZHIQUAN ⋅ Yu Mi ⋅ Hongquan Zhu ⋅ Qiufei Hu
Abstract
Benchmarks judge language-model agents mainly by whether they reach the correct final state, not by how they got there. Endpoint scoring cannot separate an agent that retrieved exactly the evidence a decision required from one that read most of the environment, nor an agent that never located the evidence from one that stopped a single action short. Settling this needs a reference cost, and authored solutions establish only that a task is solvable, not that no shorter route exists.We present DetourBench, which constrains the environment until that minimum is computable. Each task is a finite, partially observable transition system with unit-cost accesses whose resources become globally addressable after discovery. Exhaustive search yields the full-information minimum action count $L_o$ with a replayable witness; from graph-recoverable terminal states it also yields exact residual work $R_T$, which separates an episode that stopped one access short from one that wandered, while an irreversible wrong commitment is recorded as having no finite residual distance rather than assigned one. Success, attainment, residual progress, and normalized detour are reported separately. Across 300 scored episodes — five model endpoints, 20 integrated tasks, three repeats — success ranges from 63.3% to 95.0% while conditional attainment ranges from 0.798 to 0.891 and orders the models differently. Among recoverable failures, residual work separates near-complete trajectories from costly incompletion and from irreversible commitment. We open-source the instances, oracle witnesses, and evaluation harness at https://anonymous.4open.science/r/DetourBench-C955.
Chat is not available.
Successful Page Load