Steering to Persuade: A Learned Policy over an LLM's Internal Emotion Representations
Abstract
Persuasive dialogue depends not only on what an agent says, but also on how it says it. Existing systems typically control this aspect of communication indirectly, through prompts or domain-specific, predefined dialogue strategies, while leaving the model's internal representations untouched. We introduce an alternative approach that treats emotion as an actionable component of a pretrained language model's internal representation. We first establish the existence of a steerable emotional subspace in the latent representations of a frozen language model by contrasting activations across emotions for the same prompt. We then use this subspace to directly control the model's emotional expression during generation, without changing its weights. A reinforcement learning policy learns which emotional state to apply at each conversational turn to maximise persuasion. We evaluate the approach in simulated charity-donation dialogues, comparing it against prompt-based emotional control and persuasion-strategy planning. We find that latent emotional steering consistently outperforms prompt-based delivery of the same emotional targets, matches or exceeds the performance of strategy-planning methods, and transfers well to previously unseen persuasion domains.