Aligning Embodied Companion Agent Evaluation with User Experience
Abstract
Embodied companion agents, such as AI teammates in cooperative multiplayer games, converse with users while acting alongside them in shared virtual worlds. We define companion user experience (companion UX) as the user's perception of the agent as a companion, shaped jointly by dialogue, action, timing, and social appropriateness. Existing evaluation methods often assess speech and action separately, but shared-world interaction makes them mutually grounded in each other and in the evolving game state. Moreover, companion UX is difficult to specify in advance, because what feels helpful, natural, or socially appropriate varies across users and situations and often becomes visible only during play. We frame evaluation of logged or simulated interactions not as a fixed benchmark, but as an iterative alignment process that uses evidence from user tests to revise what behaviors are measured and how they are combined into judgments. In a month-long user test with 1k users and 38k sessions, we construct a paired trace--feedback corpus linking gameplay traces to UX ratings, free-text feedback, and pairwise preferences. Through this iterative alignment process, we develop an evaluation suite whose results better align with user preferences collected in user tests. A subsequently deployed model selected using the refined suite received recommendation ratings comparable to the strongest user-tested model.