NARRA-Gym for Evaluating Interactive Narrative Agents
Abstract
Interactive narrative is becoming a practical frontier for language-model agents, spanning game NPCs, collaborative story production, emotionally responsive companions, and story-grounded interface generation. Yet most current evaluations still rely on static prompts, isolated story outputs, or post-hoc ratings, leaving open whether frontier models can sustain coherent, adaptive, emotionally grounded interaction over time. We introduce NARRA-Gym, an executable evaluation environment that turns a sparse emotional seed into a complete interactive story episode and records the full agent loop: story construction, memory updates, planning, pacing control, and optional artifact generation. The benchmark jointly tests five coupled capabilities---creative story generation, long-context state tracking, character simulation, empathic personalization, and interactive artifact generation---through a fixed LLM-as-judge sweep and a human preference study with user-provided experiences. Across nine generator models and eight benchmark personas, Claude Sonnet 4.6 is the strongest and most robust performer, while Claude Opus 4.6 forms a high-variance second tier and several middle-tier models show different tradeoffs between story quality and user experience. Human rankings recover the same top tier but shift the middle of the ordering, indicating that automated judges are useful for broad screening while human preference remains important for style, tone, and emotional resonance. Finally, failure analysis shows that the key remaining bottleneck is not surface fluency but resistance-sensitive personalization: models can write polished scenes while losing track of what the user is resisting or emotionally unable to accept.