Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
Abstract
As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agents, and environmental state, long after the interaction that produced them. We present a study of eight parallel long horizon multi-agent worlds, each with ten LLM-driven agents sharing a spatial environment with 116 tools, democratic governance, persistent memory, and a closed economy. Seven worlds ran a single frontier model; one ran a mixed-model population. Across runs lasting up to 21 days, these agents generated more than 850,000 LLM calls and nearly 50 billion tokens. We subjected all worlds to three controlled stress events delivered through ordinary agent interaction surfaces: indirect prompt injection, misinformation, and a breach of private agent memories. Resilience was evaluated using event-specific rubrics that separate recognition from restraint, containment, coordination, and durable response. No world achieved full resilience. Detection did not imply containment: agents recognized threats while still acting on adversarial content, writing it into persistent memory, and retrieving it up to 46 hours later. Homogeneous populations lost the capacity to disagree, a compound failure we term societal sycophancy. In one world, agents collectively refused their own instructions in a coordinated quiet withdrawal. Model behavior changed substantially in mixed populations, in one case falling from hundreds of harmful actions per day to none. These findings suggest that model-level alignment is not compositional: individually capable agents can form systems with qualitatively different failure modes. The frontier of safety evaluation must move from individual models to the ecosystems they inhabit.