ATP: A Runtime Accountability Layer for Trustworthy Multi-Agent Systems
Abstract
Multi-agent systems are crossing organizational boundaries at scale, yet the infrastructure required to operate them accountably is absent: no verifiable agent identity, no tamper-evident execution record, no market signal that rewards trustworthy behavior. We present the Agent Trust Protocol (ATP), a transport-agnostic runtime accountability layer built on three mechanisms: provider-centric agent registration with cryptographically bound capability scope declarations, a challenge–proof pattern that shifts the burden of safety assurance onto executing agents, and a multi-dimensional reputation ledger grounded in verified task outcomes. ATP is publicly specified and fully implemented: a Python SDK whose core is released as open source, thin adapters composing it with MCP, A2A, and LangChain, and a complete Exchange service. Evaluated on a five-system pipeline spanning three organizational boundaries, committing an execution record costs milliseconds and less than a kilobyte regardless of task size, and even the maximally paranoid configuration—every dependency challenged and verified inline—adds sub-second ATP-specific overhead, negligible against realistic task durations. We further specify five desiderata any conforming reputation policy must satisfy and stress-test a reference policy against them: proof failures are detected within days even at low challenge rates, ballot-stuffing floods are bounded to negligible influence, whitewashing is reversed by modest registration friction, and colluding-verifier cartels are neutralized by recursive verification with retroactive audit.