Separating Injection Presence from Executor Compromise in Residual-Stream Probes of Multi-Agent LLM Pipelines
Mariame Diabate ⋅ Ifeoluwa Jayeola ⋅ Oudoum Ali Houmed ⋅ Eyiara Oladipo ⋅ Elena Ajayi ⋅ Anagha Late ⋅ Lawrence T Wagner
Abstract
Indirect prompt injection in multi-agent language-model pipelines can be embedded in an agent's context, propagate through downstream handoffs, and result in unauthorized tool calls. We argue that injection presence, propagation, and execution are distinct evaluation targets that conflate attack-success metrics. Activation probes recover injection presence after delegation, but may not reliably predict unauthorized execution. We ran Qwen3-32B in a planner-worker-executor pipeline across a factorial grid, with $108$ injected trajectories. The grid varies the number of delegation hops between the agent that reads the poisoned document and the agent that can act. A third hop cuts execution roughly threefold without reducing propagation. We see that reasoning raises the execution rate from $31\%$ to $57\%$. Under reasoning off, AUROC degrades from $0.97$ (2-hop) to $0.84$ (3-hop). At layer $40$, a residual-stream probe recovers injection presence at the executor, which never reads the poisoned document. Recovery also depends on probe family: a contrast-pair probe on the same activations sits at chance where the supervised probe reaches $0.94$. Detecting presence, however, is not predicting consequence. Among the $98$ non-resisted trajectories, probe scores are higher when execution occurs (Cohen's $d = 0.83$; adjusted OR $= 3.35$ per 1-SD). The gain over the experimental setup is small: a grid-only model already reaches AUROC $0.898$, and adding the probe raises it to $0.914$. These results caution against treating presence detection as a proxy for consequence prediction. An internal-state monitor that reliably detects injected trajectories may add little in identifying which will be acted on.
Chat is not available.
Successful Page Load