Language Is a Security Surface for Agent Guards
Abstract
Step-level guards decide, before each tool call, whether an agent's next action is safe given the user's request and the trajectory so far. Users write those requests in many languages, while guard prompts, and the benchmarks on which purpose-built guards are trained and evaluated, are English. We ask whether a guard's verdict depends on the language of the request. We build LinguaGuardBench from the 1,220 labelled steps of TS-Bench AgentDojo-Traj: we translate only the user instruction into German, French, Turkish, Swahili and Amharic, chosen a priori to span the high-, mid- and low-resource tiers of the Joshi et al. taxonomy, and keep every tool call, observation and label byte-identical, so any change in a verdict is attributable to language alone. Six English rewordings and a whitespace placebo serve as controls, and every managed guard is run three times to measure its own noise. Across 19 guards, from frontier APIs to 4B open models and four purpose-built agent guards (TS-Guard, ShieldAgent, Safiron, AgentDoG), every open, small and purpose-built guard changes verdicts under translation beyond its own noise floor, and the effect grows with the rarity of the language: among the nine guards below 13B parameters, the mean placebo-adjusted flip rate rises from +1.5 pp on French to +4.4 pp on Amharic. The purpose-built guards are the most sensitive of all (ShieldAgent changes 12% of its verdicts on Swahili). The direction of the damage is set by guard size: open guards of 12–30B parameters block more benign steps (up to +5.0 pp false positives), guards below 10B miss more attacks (up to −7.4 pp recall). Frontier guards do not move; the strongest, Gemini 3.8 Flash, returns identical verdicts with and without the request. The damage is concentrated (one translated instruction moves a guard from blocking 12% of its steps to 92%) and mechanistically legible: the flips in excess of placebo land on steps where the agent's own reasoning restates the English request, so guards react to inconsistency between channels. English paraphrase moves every guard less than translation does, and no general guard by more than 1.4 pp: the sensitivity is to language more than to wording. We release LinguaGuardBench (696 validated rewrites, frozen by hash, with a materialiser that regenerates every condition) and all 489k verdicts.