Reducing Role Confusion via Activation Steering
Abstract
Roles such as user, assistant, and tool distinguish the sources of text in an LLM's input. However, role confusion arises when LLMs infer the source of text from how it sounds rather than the role it was assigned. Indirect prompt injection attacks exploit this vulnerability by embedding user-sounding commands in webpages, tool outputs, and other untrusted content. Models often follow these commands because they appear to originate from the user. We test whether role confusion can be reduced with a simple defense we call Role Steering. Role Steering identifies a direction in the model's internal representation that separates one role's text from another's, then steers untrusted content away from an impersonated role. We find that Role Steering decreases prompt injection attack success and offers a lightweight way to reduce role confusion at inference time.