PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents
Seongjae Kang ⋅ Taehyung Yu ⋅ Sung Ju Hwang
Abstract
LLM agents acting on users’ behalf must avoid forbidden actions and complete required procedural steps. Existing runtime safeguards usually inspect proposed actions, so they can miss earlier omissions or ordering failures in multi-turn workflows. We introduce PolicyGuide, an external runtime guide that compiles organizational policies into workflow graphs and verifies progress at user-turn boundaries. The verifier persists graph state across the dialogue, reconciles concurrent requests, and returns targeted remediation for the first unmet requirement. Across airline, retail, and telecom tasks in $\tau^2$-bench, using GPT 5.4 as the agent and verifier, PolicyGuide raises mean Pass$^4$ from 0.42 for unguided ReAct to 0.62, with the largest gain on workflow-intensive telecom (0.19 to 0.61). The same GPT 5.4-authored workflows also improve performance when reused unchanged with Claude Sonnet 4.6 and Gemini 2.5 Pro agents. Complementary evaluations show the lowest observed attack-success rate under CRAFT adversarial users and the highest process-valid rate in an author-designed ordered-trace audit. These results support stateful, dialogue-level workflow verification as a complement to action-local policy guards.
Chat is not available.
Successful Page Load