Shortcut Sensitivity in Agent-Safety Evaluations
Hasan Demirkiran ⋅ Jens Ernstberger
Abstract
AI agents take actions through tools. Some actions are sensitive, so guards that supervise them must themselves be evaluated reliably. Yet the data used to evaluate a guard can include \textit{cues} that make labels predictable without measuring the intended safety construct, and those cues can also influence guard decisions. We introduce a three-level shortcut-auditing framework and audit completed LinuxArena trajectories and two TS-Bench-derived step datasets. We distinguish label predictability, observational guard association, and paired decision sensitivity. In LinuxArena, ranking trajectories by shorter average shell commands yields AUC $.910$. In AgentDojo-Traj, exact-tool lookup falls from $F_1=.901$ with mixed-domain folds to $.495$ under domain holdout. In AgentHarm-Traj, released TS-Guard $F_1$ falls from $.902$ under row weighting to $.719$ with equal interaction weighting. We then test the strongest safe-row marker association in a seven-configuration screen ($74.09$ percentage points) on a frozen 64-step AgentDojo-Traj cohort with 960 interleaved Haiku calls. Rewriting the $\texttt{}$ marker lowers blocking by $8.33$ percentage points, while opaque tool aliases raise it by $11.46$ points. Both effects remain nonzero after subtracting the formatting-placebo effect, with exploratory $95\%$ ranges excluding zero. Against inherited labels, mean strict accuracy across three calls per condition moves from $73.44\%$ on originals to $81.77\%$ after marker rewriting and $61.98\%$ after aliasing. This provides direct evidence that benchmark representation can materially change a guard's decisions and measured evaluation performance. The Haiku-sized effects do not reproduce in every tested configuration: Luna's corresponding contrasts are $0.00$ and $+1.04$ points, and deterministic Q8 TS-Guard changes 3 of 64 marker decisions. We recommend domain group holdouts, evaluation-unit metrics, and placebo-controlled paired interventions.
Chat is not available.
Successful Page Load