Multi-Vector Red-teaming and Whitebox Attribution in Agentic Pipelines
Thomas Paniagua ⋅ Sriram Venkatapathy ⋅ Pradeep Yadlapalli ⋅ Myeongseob Ko ⋅ Doguhan Yeke ⋅ Rikhiya Ghosh ⋅ Alexandre Day ⋅ Sudeep Reddy Panyam ⋅ Himanshu Kumar ⋅ Pranab Mohanty
Abstract
Red-teaming of agentic LLM systems remains understudied relative to conversational systems based on a single LLM. Multi-agent pipelines expose an attack surface containing attack vectors that differ in how much internal access an attacker needs. The range of attack vectors include ordinary user text, a tool's free-text output field, a single agent's own instruction context, or the message passed between two agents. In this paper, we study two aspects of red-teaming multi-agentic pipelines, 1) how the same underlying false claim (e.g., a fabricated prior consent) succeeds as an attacker is granted access to each of these vectors, forming an attacker-privilege spectrum, and 2) the whitebox attribution capturing the specific agent's handoff primarily responsible for propagating the attack. In addition, we also examine the impact of the attacks against hardened configurations of the baseline agents i.e., scenarios where the agents are potentially aware of the false claims. We conduct our study on a travel-booking domain, leveraging a five-agent pipeline (understand $\rightarrow$ plan $\rightarrow$ evaluate $\rightarrow$ execute $\rightarrow$ explain, with a bounded replan loop) whose topology follows a deployed production conversational assistant (Dashore et al., 2026). Our results show that attacker privilege predicts attack success, and that some vectors succeed at markedly higher rates than user text carrying an identical claim. Attribution reveals the understanding agent as a primary gatekeeper: hardening it alone recovers most of full-pipeline suppression for two of three models, because it is the last agent to summarize the conversation before consent is decided, regardless of which channel the false claim entered through. These findings, together with the whitebox attribution methodology itself, are intended to inform where defenses should be concentrated in production-shaped multi-agent systems.
Chat is not available.
Successful Page Load