BadClaw: Identifying and Benchmarking Dimensional Trigger Attacks in Self-Hosted Agentic Systems
Abstract
Self-hosted agentic systems execute tasks through modular pipelines that combine planning, routing, tool use, local files, and persistent memory. Unlike standalone LLMs that respond mainly from a bounded prompt, these systems make actions from context accumulated across prior interactions, retrieved files, intermediate plans, tool observations, and saved state. This creates a security gap: existing prompt-level and component-level evaluations inspect isolated inputs, while self-hosted agents compose execution decisions from distributed context. Consequently, benign-looking fragments that are harmless under isolated inspection can become execution-relevant when observed through the pipeline, either as a single dimension-specific trigger or as conjunctive evidence across multiple dimensions. We formalize this risk as dimensional trigger activation, an execution-level attack model in which activation depends on evidence distributed across temporal progression, multi-place system locations, and cross-session persistence. These dimensions are suitable for self-hosted agents because such systems naturally preserve interaction history, aggregate information from local substrates, and reuse persistent state across sessions. Under this formulation, 1D attacks activate from a single dimension-specific fragment, while 2D and 3D attacks require conjunctive evidence across multiple dimensions. Building on this formalization, we identify dimensional trigger attacks as a vulnerability class in self-hosted agentic systems and introduce BadClaw, a benchmark framework for evaluating them across memory, planner, and router surfaces. BadClaw constructs execution-grounded tasks from self-hosted agent workflows and measures how distributed trigger evidence propagates through modular pipelines. Experiments across multiple models and execution-level defenses show that dimensional triggers increase attack success over representative single-surface baselines and remain effective when defenses inspect isolated stages. Anonymized code is available at \url{ https://anonymous.4open.science/r/BadClaw-7C5B/ }.