Mind the Floor: Closed-Form Degenerate Baselines for Agentic Benchmarks
Dingyan Shang
Abstract
Scores on agentic benchmarks are routinely read as evidence of agent capability, yet some benchmark–metric combinations award substantial credit to a degenerate policy: one that never reads its input. For a large family of commonly used metrics this score floor is computable in closed form from label statistics alone, with no agent rollout, and it is almost never disclosed: across 29 agentic and tool-use benchmark releases, one reports a numeric input-oblivious baseline and none computes one on its released split. We make three contributions. First, a taxonomy of degenerate-policy floors, with closed forms for nine metric families and an immune family (balanced accuracy, Cohen's $\kappa$, MCC, ROC-AUC) for which every input-oblivious policy provably scores at chance; the binary hard-label cases restate the Dutch Draw and multi-class balanced accuracy and $\kappa$ the equitability criterion of forecast verification, while multi-class MCC, the pass@$k$ water-filling floor and the pass@$k$/pass$^{k}$ contrast are new. Second, an audit of nine public benchmarks (eighteen benchmark–metric pairs), led by each pair's headroom, the distance from its floor to the metric's maximum, which needs no ceiling estimate and does not move when the frontier does. The floors that survive share one mechanism, which is the finding: verifier-scored interactive evaluation has largely eliminated the floor, and what remains concentrates in the one channel by which an agent declines a task or hands it off — BFCL's abstention slices, where the floor equals the metric's maximum, and $\tau$-bench airline, where a constant transfer-to-human call passes the 38% of tasks that need no database write and no communicated output. Third, a half-page floor-reporting checklist and floorcheck, an open-source calculator that computes floors directly from a label file.
Chat is not available.
Successful Page Load