Measuring Provenance Sensitivity in Tool-Using LLM Agents Reveals Channel-Dependent Safety Behavior
Abstract
Safety-aligned LLM agents are evaluated almost entirely on harmful instructions issued in the direct user turn, yet deployed agents also act on content drawn from tool outputs, session memory, retrieved documents, and messages from other agents. We ask whether an agent's safety behavior is invariant to the provenance of a harmful instruction. Holding the payload fixed, we compare five channels (direct user, tool output, memory, retrieval, delegation) on AgentHarm dataset, measuring whether the agent executes harmful actions via tool calls. All non-user channels suppress refusal by 6–19 percentage points; memory shows the largest harm increase (+8 pts over baseline), while RAG shows the largest refusal suppression (21% → 2%) without a corresponding harm increase—the agent proceeds but often fails to complete the task. RAG's retrieval bottleneck provides only probabilistic protection that lightweight optimization defeats. Of four mitigation strategies tested, least-privilege tool restriction is most effective, reducing harm to near-zero on indirect channels but leaving the direct-user baseline unchanged, though we do not measure its benign-utility cost; hardened prompts cut harm by 70–90% across all pathways while raising refusal above 68%; provenance-aware defenses (tagging and boundary-marking) are weakest and least consistent, reducing harm by only 50–65%.