When Agents Go Rogue: Provably Irreversible Shutdown
Abstract
Autonomous agents may drift from their intended behavior, be subverted by an attacker, or actively turn against their own creator. When that happens, the creator needs a reliable, irreversible way to stop the agent that works even after a party controlling the agent's host has become adversarial. We formalize a cryptographic mechanism that gives the creator exactly this: a secret stop-signal, known only to the creator and never learned by the agent, whose presentation forces the agent's immediate, irreversible disablement with no newly authorized action ever honored, even against an adversary who controls the agent's own machine and can copy, edit, snapshot, or rewind its software. Since destruction cannot be enforced in adversary-controlled software, enforcement rests on a minimal trusted, non-restorable device that gates all actions, so the agent cannot act without passing through it.