Entropy Signatures of Adversarial Suffixes for Indirect Prompt Injection in Tool-Using LLM Agents
Abstract
Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection: an adversarial instruction hidden in a tool observation. We study defended cases where the bare injection fails but a fluent, optimization-based adversarial suffix induces an attacker-chosen tool call while evading a standard perplexity filter. We ask whether predictive entropy provides another simple, training-free statistical signal before the agent acts. We test this with CPD Online, an entropy change-point monitor, on untrusted tool observations across five attack-generation methods against Llama-3.1-8B, Qwen3-4B, and Ministral-3-8B. On fluent AdvPrompter attacks, adding CPD Online to a held-out mean- and windowed-perplexity screen raises recall by 21 and 44 percentage points on Llama3.1 and Ministral3 under the same 5% training-benign false-positive constraint. Detector-blind TAP attacks confirm transfer beyond AdvPrompter: against 1,208 benign outputs, CPD Online reaches 0.99, 0.93, and 0.97 AUROC on Llama3.1, Qwen3, and Ministral3, and exceeds every evaluated windowed-perplexity configuration on Llama3.1 and Ministral3. Among IterInject attacks passing every evaluated perplexity threshold, CPD Online detects 90% on Qwen3 and 80% on Ministral3. Perplexity filters remain effective when suffixes contain locally improbable tokens; CPD Online can instead detect fluent suffixes whose anomaly appears as a sustained predictive-entropy shift. Finally, form-and-length-matched controls and likelihood-adaptive attacks reveal a boundary shared by these statistical monitors: they detect attack-induced distribution shifts, not malicious intent. We release the five-family corpus, code, and evaluation protocol.