Who Controls the Collective AI Vote? Real-World Evidence and Adversarial Testing of Multi-Agent Delegation in DAO Governance
Abstract
AI systems increasingly recommend and execute group decisions, yet several visible agents may still share one actor who controls their instructions, evidence, software, signing, or reversal. Organizational authority, collective choice, and AI reliability study parts of this problem but rarely the full delegated decision. We ask: when AI agents recommend, aggregate, and execute a collective vote, who can change the result, and which safeguards preserve independent judgment and effective correction? Decentralized autonomous organizations (DAOs) make authority observable through public proposals, recommendations, rules, and actions. We combine an analytical model, a public-record audit of three deployed systems, exhaustive tests of three voting rules, one-agent interventions, a paired open-model replay, and a synthetic stress test. Multiple agents did not necessarily imply multiple controllers, and the same recommendations led to action or abstention under different rules. For one fixed model, separating untrusted proposal text from governing instructions reduced targeted vote changes from 32 to 4 among 39 comparable responses. In the simulator, controller separation reduced common-input attacks but not direct one-agent compromise; layered safeguards lowered all four channels. The paper connects residual authority, rule semantics, and reliable-agent engineering, giving communities, developers, and auditors a practical way to locate control, test failure paths, and state what their evidence does—and does not—support.