AGovBench: A Benchmark for Evaluating Agentic Governance in Consequential Decision Systems
Abstract
As LLM agents increasingly participate in consequential multi-stakeholder decisions, evaluation must capture not only whether they coordinate, but how their interventions affect the systems they govern. Existing benchmarks rarely connect deliberation, implemented actions, and independently measured downstream consequences. We introduce AGovBench, a framework for evaluating agentic governance through linked process, intervention, and outcome measurements. In a consequential hiring instantiation, role-conditioned agents negotiate bounded modifications to an editable causal decision system; fairness, predictive performance, and stakeholder effects are recomputed from resulting states, with decision stability measured across episodes. Across 47 configurations and 235 governance episodes, consensus was common but frequently coexisted with fairness conflicts, predictive-performance losses, and stakeholder harm. Governance behavior also varied across models and was sensitive to evaluation feedback, interaction ordering, semantic representation, and reasoning configuration. Increased reasoning produced opposing policy shifts across evaluated model families without uniformly improving outcomes. AGovBench therefore treats coordination as an auditable process rather than a sufficient criterion for successful governance.