Runtime Verification of Multiple Natural Language Criteria for Agent Governance
Silviu Pitis ⋅ Parand A. Alamdari ⋅ Jessica Tang ⋅ Toryn Klassen ⋅ Sheila McIlraith
Abstract
AI agents are governed by hundreds of natural language criteria---yet existing verifiers cannot efficiently evaluate an agent's output against a large set of such criteria and report which are satisfied, violated, or inapplicable. We introduce the VFM, a tree-attention verifier that encodes the shared context and target text, and scores $N$ natural language criteria in parallel, returning independent per-criterion ternary verdicts. The VFM is pretrained on CriteriaBank, an open corpus of 357K (context, target, criterion, verdict) tuples spanning over 285K distinct criterion strings. A finetuned VFM-4B reaches 88.3\% accuracy on a synthetic governance benchmark, surpassing larger generative judges, and on real insurance-compliance calls it is competitive with GPT-5.4 with synthetic finetuning alone. On the DynaBench multi-criterion benchmark, a finetuned VFM-4B reaches $0.85$ F1 on trace-level failure detection versus DynaGuard-4B's $0.72$; the VFM additionally localizes the violated rule as its top-1 prediction in $98.4\%$ of failing traces, and a top-3 cascade to a 26B Gemma 4 chain-of-thought judge reaches $0.94$ F1. CriteriaBank pretraining improves cross-domain transfer on four additional benchmarks, particularly in the low-data regime. On an H100, VFM-4B evaluates 1000 criteria against a 4K-token context in under 2s at $<$10GB VRAM, making it suitable for runtime verification. We will release CriteriaBank, trained checkpoints, and evaluation code.
Chat is not available.
Successful Page Load