Verifying Persistent Constraint Compliance in Multi-Turn LLM Agents
Abstract
Verifying a multi-turn agent means answering whether it still obeys the constraints its user stated earlier. We study a verification recipe that blends a deterministic, environment-grounded check with an LLM judge and reports an aggregate constraint-drift rate alongside it. We audit that recipe in MACS (Multi-Agent Commerce System), a fixed-catalog conversational recommendation environment in which catalog-expressible constraints admit deterministic SQL checks, letting us audit both the strengths and the coverage limits of environment-grounded verification. Each part of the recipe can stop informing while its reported value still looks healthy. First, saturation: on our 10-scenario, K=5 benchmark the deterministic component sits at its ceiling for 99% of runs (mean 1.000, SD 0.000 for two of three systems), so despite a nominal weight of 0.55 it contributes 1.5% of the composite's variance, and passing reduces to an LLM-judge threshold. Second, coverage confounding: aggregate drift is computed only over turns where a system returned products, and our baselines do so on 65% and 44% of constrained turns against 100% for MACS, so their near-zero drift partly rewards abstention. Third, blend masking: a single composite hides which component moved. We then show a turn-level, opportunity-adjusted verifier that restores discrimination by scoring every encoded machine-checkable constraint on each checkable turn, reporting coverage as an explicit denominator, and separating checks requiring cross-turn retention from checks answerable from the current turn. Under matched denominators it isolates a controlled ablation that the aggregate signals blur: removing the agent's persistent constraint store leaves all 60 just-stated checks satisfied while violating 10 of 55 carry-over checks (95% CI [0.00, 0.55]; concentrated in one of ten scenarios, where it recurs in 5/5 runs). Decomposing the same ablation separates two failure channels the blended score conflates: one the checker catches (1.00→0.34), and one it scores a perfect 1.00 while the agent silently drops a brand the user asked for in all 5 runs, because that constraint is not in the ground truth. Only the LLM judge caught it. Auditing the benchmark, 7 of 10 scenarios state a constraint no deterministic path reads, and 7 of those 9 constraints have no catalog field behind them at all, which bounds what environment-grounded verification can cover and partly explains the saturation. We also report a verifier defect found in our own harness, in which a silently defaulting judge produced an implausible 100% pass rate for baselines. We argue that agent verification signals should be reported with their denominators, their ceilings, and their per-component movement, since none of these failures is visible in the aggregate.