SPILLWAY: Separating Authorization from Sanitization to Defend Tool-Using Agents against Prompt Injection
Woosang Lim ⋅ Yeseul Chang ⋅ Jaewoo Lee
Abstract
A tool-using agent faces two trust boundaries that fail for opposite reasons: what it reads from untrusted tool output and what it externalizes as an effect. Existing defenses place security machinery at different points, including sanitization, capability or policy enforcement, and model-mediated planning or validation. SPILLWAY assigns the two boundaries different mechanisms. Its target-authorization floor is a deterministic egress membership test derived from schema-declared identity roles; a separate clause-level ingress sanitizer redacts clear injected imperatives, keeps bare data, and escalates only ambiguous clauses to a small local model. On AgentDojo, SPILLWAY records zero observed executed-effect in-scope ASR across twenty backbone-by-suite cells, while its unmodified raw AgentDojo ASR is $.008$, the lowest of any method on the benchmark's own scorer (next best $.028$). It retains $97.9\%$ of the undefended agent's under-attack mean utility, exceeding the undefended agent outright on two of four suites; MELON and Progent retain $64\%$ and $84\%$ and have in-scope leaks on twelve and thirteen cells. Two static defense-aware stress tests also yield zero observed in-scope ASR, including chat-template forging handled by a deterministic turn scrubber. The contribution is the composition: authorization stays a schema-derived lookup, and model compute is spent only on residual ingress ambiguity, never on the authorization floor. This makes SPILLWAY both effective and efficient, holding in-scope attack success at zero while adding only a fraction of the runtime cost of defenses that re-run the agent.
Chat is not available.
Successful Page Load