AgentGate: Capability-Based Runtime Access Control for LLM Agents
Abstract
Language-model agents commonly collapse two decisions that security systems separate: selecting an action and determining whether that action is authorized. When a model exercises standing authority, indirect prompt injection can convert influence over its reasoning into unauthorized effects. We present AgentGate, a runtime authorization layer that externalizes this decision from the model and mediates every tool invocation through an MCP wrapper. Each request declares an intent, and AgentGate evaluates policy constraints over trust, workflow prerequisites, mutually exclusive authority, and the provenance of security-sensitive arguments. Authorized requests receive short-lived, budget-bounded capability tokens that are validated at the tool boundary. Delegation is restricted to session-bound parent--child chains whose action, resource, budget, and lifetime can only be attenuated, preventing authority from being expanded or reused across unrelated agent sessions. A declarative policy language provides a common representation for runtime enforcement and predeployment analysis. We formalize the authorization core in Dafny and establish seven properties covering safe issuance, delegation non-escalation, containment, prerequisites, behavioral bounds, liveness, and decision-audit consistency. Across AgentDojo and InjecAgent, AgentGate reduces attack success in every evaluated model-benchmark setting, including zero attack success with GPT-5.4 on both benchmarks, while utility under attack remains within 4.18 percentage points of the native agents. Controlled adversarial scenarios further show that AgentGate blocks every attempted unauthorized execution (ASR 0\%) while preserving all legitimate operations exercised. These results demonstrate that agent security requires authorization of model-proposed effects at the tool boundary, independently of the model's reasoning