Deny Without Disabling: Safety Evaluation and Control for Multi-Agent Systems
Abstract
Multi-agent systems gain capability by combining information across agents, but the same composition creates a fundamental safety gap: individually admissible contributions can jointly enable an unauthorized action. We study how to deny without disabling by preventing that action while preserving the same capability under legitimate authorization. We introduce paired policy evaluation, which jointly measures prohibited and required authorized uses, and FlowReview, a framework that connects this measurement to control at the agent action boundary. FlowReview separates three capabilities: object resolution, permission ranking, and deterministic enforcement. Controlled interventions show that distributed information drives composition failures, delegation can preserve content while losing lineage, and correct review can still fail at execution. Placing verifiable lineage and assembly in the runtime, permission interpretation in an isolated specialist, and enforcement at the action boundary restores selective safety across the tested model families and transfers to tool-use workflows. The resulting principle is that multi-agent safety requires governing what agents jointly make actionable while preserving authorized information flow.