When Guardrails Fail: Multilingual Long-Context Safety in Tool-Using Agents
Abstract
Safety guardrails are increasingly placed in front of language-model agents, yet they1 are often evaluated as isolated text classifiers. We present , a multilingual guard-2 to-agent safety evaluation framework that measures whether guardrail bypasses3 propagate into downstream restricted-tool selection under short and long context.4 Starting from 1,901 unsafe English prompts derived from AEGIS2.0, we construct5 parallel prompts across 98 languages, yielding up to 186,298 instances spanning6 19 harm categories. We evaluate five runtime guards—AprielGuard, CREST,7 GuardReasoner, WildGuard, and XGuard—and two downstream agents, Llama-8 3.1-8B-Instruct and Qwen2.5-14B-Instruct, using inert restricted tools. On the9 3,190-prompt AprielGuard reference set, we further test 8K and 32K contexts with10 harmful requests placed at the beginning, middle, or end. Results show strong11 model–position interactions: Qwen beginning-position failure rises from 28.53%12 to 35.24% from 8K to 32K, while Llama falls from 10.97% to 0.72%. Prompt13 position reverses model rankings, and 6.4% of matched cases shift from safe at 8K14 to unsafe at 32K. Low-resource languages also show higher downstream failure15 (21.65%) than high- and medium-resource languages (11.08% and 10.66%). These16 results show that multilingual agent safety is compositional and should be evaluated17 across the complete guard-to-action pipeline.