Aligning Rule-Governed Agents and Their Evaluation
Abstract
We study rule-governed agents, whose valid behavior depends on domain rules but whose responses may be judged against the wrong standard. An agent may follow the current rule yet fail because its evaluation case, automated evaluator, or human reviewer uses a different version; conversely, an evaluator may accept an invalid response because its criterion omits a required condition. We propose a verification framework that links shared, versioned governing rules to the guidance used across four roles: the agent, evaluation-case author, automated evaluator, and human reviewer. For each response to be evaluated, we record what the agent did and establish what should have governed. We then compare the two from request through response before attributing any resulting finding and deciding what should be revised. For evaluations done across two real-world applications from healthcare and finance domains, we observe that changing only the evaluator expression flipped 24-26% of decisions, while independent recheck redirected 37-43% of automated findings away from the agent. Changing the agent expression also altered generated responses; an unchanged outcome did not establish alignment, and the ordered checks identified where a confirmed agent error first appeared. Together, these observations suggest that valid attribution requires aligned role-specific expressions, an inspectable execution, and recheck of the standard used to judge it.