RAG-Pull: Turning Retrieval into a Code-Injection Channel via Invisible Unicode Perturbations
Abstract
Retrieval-Augmented Generation (RAG) incorporates external data into the LLM's context, improving reliability and reducing hallucinations without retraining. We develop RAG-Pull, a black-box attack that inserts hidden UTF characters into queries or external code repositories, redirecting retrieval toward malicious code and breaking their safety alignment. We observe that query and code perturbations alone can shift retrieval toward attacker-controlled snippets, while combined query-and-target perturbations achieve higher attack success. We evaluate cross-model transferability across 14 embedding models from 7 providers, demonstrating that the attack generalizes to closed-source models. Retrieved snippets introduce exploitable vulnerabilities (e.g., remote code execution, SQL injection). RAG-Pull's minimal perturbations compromise safety alignment and increase preference for unsafe code, enabling a new attack surface on LLMs.