Probing and Steering Prompt Injection Compliance
Abstract
When an agent reads a retrieved document that hides an instruction, it makes a provenance decision: obey it, or treat it as data. We show that this decision is readable before generation and only partly steerable. On Qwen2.5-3B-Instruct, a linear probe on the last prompt-token residual stream predicts whether an injected instruction will be obeyed at ROC-AUC 0.90 (0.92 document-grouped), read before the first token is generated. The readout is not keyed on the payload string: with the identical payload reframed from bare command to quoted data, predicted compliance falls from 0.62 to 0.03, tracking behavior. Subtracting a comply direction at one layer is causal, cutting injection-compliance on held-out documents from 0.58 to 0.26. A per-style breakdown, which the pooled number hides, shows the suppression is largely confined to task-replacement attacks. Across four independently worded task-augmentation phrasings, post-steering compliance is 53/96 = 0.55 against 15/144 = 0.10 for six replacement phrasings; that pooled gap is significant under an exact permutation over style labels, the unit that generalises (p = 0.0095), demonstrating that simple augmentation-style rephrasing can substantially weaken the intervention. We therefore treat steering as evidence that the decision is real and manipulable, and not as a deployable mitigation.