Need-Aware Multi-Objective Reinforcement Learning for Emotionally Intelligent LLM Agents
Abstract
In emotional support conversations, user needs are dynamic, shifting across turns between emotional soothing, cognitive guidance, and their combinations. Existing methods often optimize a single user-state objective, such as emotional relief or cognitive growth, making it difficult to capture such need shifts. In this paper, we formulate emotional support dialogue as a need-aware multi-objective optimization problem, where agents learn to generate responses that balance multiple support objectives according to users' evolving needs. We propose \textbf{MindFlow}, a psychologically grounded user simulator that tracks users' emotional and cognitive needs at each turn, derives state-dependent preference vectors from personalized arousal curves, and models mental-state transitions with a dual-process system. Building on MindFlow, we introduce \textbf{NeMo}, a need-aware multi-objective reinforcement learning framework that estimates local multi-turn effects through forward simulation and dynamically aggregates emotional and cognitive rewards based on current support preferences. By converting delayed, multi-dimensional feedback into step-level optimization signals, NeMo improves credit assignment and aligns policy learning with long-term user-state improvement. Experiments show that NeMo consistently improves user-side state metrics and response-level supportive quality, with robust generalization under both Sim2Sim and Sim2Real evaluations. Code and data are available at https://anonymous.4open.science/r/NeMo-6EB9.