Human Before the Loop: Constraining the Action Space Improves Security, Reliability, and Capability in AI Agents
Abstract
There has been a rise in the use of multipurpose Artificial Intelligence (AI) agents, as well as a rise in data breaches caused by these same AI agents. These types of agents are prone to an attack called prompt injection, which can cause a large language model (LLM) to be instructed to find and deliver sensitive data to the attacker. Every proposed defense carries its own trade-off. Post-training a model to refuse malicious instructions cannot anticipate attacks that have not been seen yet. Runtime supervision slows operation. Security middleware only stops the attacks it was built to recognize. Defenses that restructure an agent's control flow have documented drops in benign task completion. To solve this, we present Human Before the Loop, an agent framework where the action space and data access are determined in the design phase upfront, leaving no runtime path open to sensitive data or destructive actions. Rather than relying on traditional open-ended tool calling, our framework uses a method we term semantic action abstraction and navigates actions within a Hierarchical Extended Finite State Machine (HEFSM). This eliminates the chance of data exfiltration while increasing benign task success and realized capability. In our evaluations, our framework prevented all prompt injection attacks, whereas the baseline tool-calling agent failed to prevent 8.5% of attacks and delivered sensitive data to the attacker. Benign task completion rose from 86.7% with the tool-calling agent to 92.4% with our framework. Our framework also enabled a 30B parameter model to complete the Pokémon Red tutorial phase in 341 turns, while the same model using a tool-calling harness made no progress beyond the initial stage across 1,000 model calls. In our framework that 30B model finished the tutorial in 22% fewer model calls than the larger model it was distilled from. This method supports the creation of autonomous AI agents that can be trusted to run fully unsupervised without a trade-off between security and capability.