PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents
Abstract
LLM agents act on behalf of organizations through tool calls but do not reliably follow company policies embedded in their system prompts. Policy adherence is a dialogue-grounded safeguarding problem: failures can arise from adversarial attempts to bypass rules or from honest workflows that skip identity checks, prerequisite reads, offers, or confirmations. We introduce PolicyGuard, a sub-agent verifier that intercepts mutating tool calls, reads the full interaction history, and evaluates pending actions against raw policy text and an LLM-generated per-tool checklist. It either authorizes the call or returns conversation-specific remediation. On tau2-bench airline, paired PolicyGuard verifiers improve Pass⁴ by 12, 6, and 12 percentage points for GPT-5.4, Claude Sonnet 4.6, and Gemini 2.5 Pro, respectively; combined paired tests support overall and policy-violation gains. On the full retail and telecom base splits, PolicyGuard produces a one-task overall Pass⁴ gain in each domain without reducing mutation success. Removing dialogue from the verifier collapses mutation success, while fixed-payload adversarial probes retain positive margins over the baseline.