Algorithmic Taint Tracking: Attention-Based Provenance Enforcement for Secure Tool-Augmented LLMs
Akshat Sharda ⋅ Zhenyu Zhang
Abstract
Tool-augmented LLMs that invoke privileged operations on untrusted context are structurally vulnerable to the confused deputy problem. We introduce Algorithmic Taint Tracking (ATT), a framework requiring no weight modification that tracks token-level provenance via a parallel sideband tensor propagated through the model's own attention weights. ATT decouples detection (attention-mediated taint propagation) from policy (modular logit gating that blocks privileged token generation when taint exceeds a calibrated threshold). On single-turn tool invocations across two model families at two scales (Llama 1B/8B, Gemma 2B/9B), ATT achieves 98.1-100\% TPR with 0-5.1\% FPR vs. 0.6-43.2\% TPR for perplexity filtering, 78.3\% for RTBAS, and 0-90.1% for LLM-as-Judge on stealth attacks at ~18% inference overhead. Extended to multi-turn settings via cross-turn taint decay, ATT maintains 98-100\% TPR with 0-4\% FPR on delayed-invocation AgentDojo traces. We introduce a multilingual stealth semantic coercion benchmark (355 scenarios $\times$ 7 languages) using only benign business vocabulary, revealing a helpfulness trap: more capable models are more vulnerable to implicit coercion. Evaluation of five frontier commercial models (GPT-5.5, Claude Opus 4.8, etc.) confirms that scaling alone does not eliminate this vulnerability. ATT resolves it because it is agnostic to payload semantics, and ablation across four structurally different propagation rules confirms the underlying detection signal is robust to formulation choices.
Chat is not available.
Successful Page Load