Runtime Verification of Large Language Model Agents Under Temporal Input/Output Logic
Abhishek N Kulkarni ⋅ Bastien Bernath ⋅ Anuj Tambwekar ⋅ Vele Tosevski ⋅ Aditya Karan ⋅ Leif Hancox-Li ⋅ Tim G. J. Rudner
Abstract
Scalable, reliable, and robust monitoring of an AI agent's compliance with operational policies is central to AI governance. These policies often encode temporal normative constraints, including conditional and contrary-to-duty (CTD) norms that span multiple turns. LLM judges---widely used for evaluating compliance---are unreliable because they lack robust temporal reasoning, may degrade over long contexts, and are prone to conflating grounding with reasoning errors. We introduce a two-stage neurosymbolic pipeline for monitoring compliance of AI agents with norms specified in a novel temporal I/O logic. The first stage uses an LLM to ground each utterance into atomic propositions. The second stage uses Runtime Norm Compliance Monitor (RNCM), a deterministic, automata-theoretic monitor synthesized from temporal I/O specifications to determine compliance. We formalize the syntax of temporal I/O logic and define a graded compliance semantics that quantifies progress along reparation CTD chains. We prove soundness, completeness, and anytime correctness of the procedure, and show that RNCM scales linearly with number of norms. We evaluate the pipeline on $\tau$-bench retail conversations and find that our method significantly outperformes six open- and closed-weight frontier AI LLM-as-judge baselines. Ablations show that all errors arise from proposition grounding, empirically validating the separation of grounding from reasoning.
Chat is not available.
Successful Page Load