RedInject: A Self-Evolving Harness for Indirect Prompt-Injection Evaluation in Agentic Gyms
Tharindu Kumarage ⋅ Qian Hu ⋅ Rahul Gupta ⋅ Charith Peris
Abstract
As tool-using agents are deployed across diverse domains, indirect prompt injections, where adversarial instructions embedded in environment content redirect an agent during a benign task, pose a growing safety risk that evaluations must keep pace with. Yet most existing evaluations author safety scenarios from scratch, often within a single environment or a small set of safety-specific tasks, limiting both coverage and workflow realism. We observe that utility gyms, environments built to test agent competence, already encode domain tasks, tools, policy constraints, state transitions, and scoring that characterize realistic deployments. Repurposing them for safety evaluation would ground attacks in genuine workflows and scale testing to any domain an existing gym covers. However, the transformation is non-trivial: attacks must be inserted without changing the benign objective, eliminating valid solutions, or invalidating the original utility scoring. We present RedInject, an agentic framework that automatically compiles heterogeneous utility gyms into indirect prompt-injection evaluations at scale through automated gym structure detection and taxonomically guided attack insertion, while preserving each gym's original utility verifier. We evaluate RedInject on $\tau^2$-Bench and Terminal-Bench with eleven LLMs, finding substantial attack success and clear model separation, indicating that the generated tasks are promising candidates for safety evaluations.
Chat is not available.
Successful Page Load