Hardening Agentic Web Populations Through Multi-Criterion Competitive Co-Evolution: Lessons from Wargame Strategy Optimization
Abstract
Agents on the open web will face adversaries that keep adapting, so a defense that is only tested once will eventually be found and broken. We propose treating jailbreak attacks and safety guardrails as two populations that evolve against each other, rather than as a fixed attacker probing a fixed defense. This idea comes directly from wargame strategy optimization, where two-population competitive co-evolution with alternating evolution cycles and cross-play evaluation has already been shown to produce strategies that are robust to a moving opponent, and where a companion line of work has shown that the same machinery can measure how much competitive pressure a population must absorb. However, real-world applications can never rely on a single criterion as a key performance indicator, rather must use multiple conflicting criteria, such as, cost and efficiency. We describe how this mechanism transfers to language model safety, argue why the resulting multi-criteria, competitive and coevolutionary arms race is a genuinely networked, population-level problem rather than a two-agent toy setting, and lay out the open problems that stand between the wargame version of this idea and an agentic web version of it.