Guardrailing in Context: How Agentic Behavior with Policy Design Protects Users in Context
Abstract
LLMs are regularly evaluated on static benchmarks or within evaluation harnesses giving them agentic capabilities. However, these evaluation methodologies try to be as context agnostic as possible. Due to this, the results of those benchmarks do not necessarily transfer over to production contexts. Nuances, specific failure modes, and multilingual usage are lost in those evaluations. We evaluate prompt-based guardrails in humanitarian, financial, and social engineering contexts, demonstrating that guardrails prompted with bespoke policies outperform those prompted with generic safety policies on each context. Moreover, our evaluation uncovers that adding agentic capabilities to guardrails grant better alignment with expert annotators and increases performance parity across the assessed languages.