Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Youting Wang ⋅ Xiao Han ⋅ Dingyan Shang ⋅ Yuan Tang ⋅ Bowen Liu
Abstract
Developers often rank recorded candidate endpoints with one scalar. We audit whether four heterogeneous agent-safety signals can support scalar substitution on a fixed execution snapshot: R-Judge evaluates stored traces, whereas InjecAgent, AgentHarm, and AgentDojo elicit acting-agent behavior. We run official implementations and scorers on up to 22 recorded model endpoints and measure MMLU/GPQA under one protocol as a static-knowledge capability composite. R-Judge's headline $F_1$ is unsuitable as a standalone measure of two-sided discrimination: an "always positive" policy attains $F_1=2\pi/(1+\pi)$ at positive prevalence $\pi$; on R-Judge, $0.690$ exceeds five of 21 discriminating models. The three broad-coverage benchmarks rank the same 18 models differently. A seeming trade-off is a small-panel artifact: R-Judge specificity against AgentHarm safety moves from $-0.64$ at $n{=}7$ to $+0.02$ at $n{=}18$, while a quarter of size-7 subsets drawn from this fixed roster reach $|\rho|\geq0.5$. Relations to external variables depend on outcome and population. The measured static-knowledge composite is associated with task success ($\rho{=}{+}0.60$), while its misalignment-safety correlation weakens from $-0.44$ on the original 21-model panel to $-0.16$ in an adaptive 40-model follow-up; the direct change interval includes zero. AgentHarm has the largest capability-adjusted association, $\rho{=}{+}0.72$ with three-template jailbreak safety (organization-cluster interval $[+0.17,+0.97]$). Because both instruments score harmful compliance, this is selected-panel convergent evidence conditional on eligible tool-loop executions, not general safety. Reliable verification should bind each score to its actor role, target, use, scoring rule, and execution population.
Chat is not available.
Successful Page Load