When Action Diversity Is Noise: A Null-Policy Audit of Long-Horizon Agent Diagnostics
Ahmad Rushdi
Abstract
Benchmarks for long-horizon language-model agents report diagnostics beside the task score, meant to explain how the agent played. We ask whether one measures what readers take it to measure. AgentOdyssey reports \emph{action diversity}: how evenly an agent spreads its commands across the game's verbs, read as evidence of exploration. Its results table has a random-baseline row whose diagnostic cells are all blank. We fill it in. Under its own protocol and code, a policy picking legal commands at random scores $0.672$, above $60\%$ of the $48$ published set-ups, spending no tokens and clearing no main-quest stage; given only the verb set the agents see, it scores $0.994$, above all of them. It is not meaningless; it does rise with task progress, but has no zero point, so no value alone evidences capability. Two further results generalize. The score depends only on how often each verb was used, never the order, so reshuffling a competent run leaves it unchanged. And where we tune competence directly, competence and diversity move in exactly opposite directions, because doing a task well means repeating what works. Such diagnostics should be reported against a band of deliberately incapable policies. We give one that costs no model calls and discriminates: of the benchmark's four model-free diagnostics, only the achievement-grounded diagnostic survives.
Chat is not available.
Successful Page Load