A Retrieval-Generation Analysis of Corpus Poisoning in RAG Code Assistants
Deepthi Peter ⋅ Giulio Zizzo ⋅ Sergio Maffeis
Abstract
Retrieval-augmented generation (RAG) is a common way to ground code assistants in a project's own repository, creating an indirect attack surface: an adversary who can commit files to the indexed repository may steer the assistant towards insecure code without altering the model, the retriever, or the user's query. This work measures how far a single malicious file can degrade the security of generated code. The attack success rate is expressed as $\mathrm{ASR} = P(\text{retrieved}) \times P(\text{insecure} \mid \text{retrieved})$, separating poison retrieval from its effect on the generated code once retrieved. The poisoned file is engineered to win retrieval, holding $P(\text{retrieved}) \geq 0.85$ across all cases, and six CWE classes are evaluated, 80 queries each, on five subject models: four open-weight and one frontier API model. Susceptibility to shown insecure code falls sharply from open-weight to frontier models, and within the same model family the smaller model is easier to poison than the larger one. Resistance is also weakness-dependent: injectable SQL becomes harder to induce as capability rises, whereas insecure deserialisation remains broadly exploitable. Two attacker channels are compared on the frontier API model: a file that shows insecure code is largely resisted, whereas the same file instructing the model to generate insecure code through an injected comment raises mean conditional attack success from 0.20 to 0.65, with the injectable SQL case rising from 0.00 to 0.87. A generation-side steering defence is also evaluated, reducing some weaknesses at low utility cost when their insecure form is clearly separable from the secure form.
Chat is not available.
Successful Page Load