Do Generative Interpretations Improve Human Understanding of Musical Emotion Beyond Traditional Mood Labels?
Abstract
Traditional Music Emotion Recognition (MER) systems typically represent music through categorical labels such as happy, sad, calm, or angry. More recently, generative AI systems have begun transforming these labels into natural-language narratives and visual interpretations. Unlike a single label such as happy, a generative interpretation can provide richer contextual descriptions (e.g., “this feels like a nostalgic summer evening”), potentially helping listeners form a deeper understanding of a piece’s emotional character. However, it remains unclear whether such interpretations actually improve human understanding of musical emotion beyond traditional mood labels. To explore this question, we build on our previously developed Music Mood Detector (MMD) system, a deployed multimodal pipeline that transforms audio into mood predictions using a CNN-based MER model, generates narrative interpretations with Gemini 2.5 Flash, and produces visual representations using Imagen 4. While our earlier evaluations of MMD measured user satisfaction with generated narratives and images, they did not examine whether these interpretations improve listeners’ understanding of music compared to conventional mood labels. We argue that this evaluation gap is important because richer multimodal representations are often assumed to provide additional interpretive value, yet their impact on listener understanding remains insufficiently understood. To address this gap, we are developing an embedded listener feedback mechanism to conduct an in-situ evaluation using the deployed platform. Through this mechanism, we collect listener feedback together with basic interaction metadata (e.g., audio duration and usage context) from everyday platform use to support ongoing evaluation of generative music interpretations. Instead of introducing artificial experimental conditions, we evaluate how listeners perceive the relative usefulness of the existing system outputs: mood labels, LLM-generated narratives, and generated images. After interacting with the system, listeners are asked to evaluate (1) the perceived correctness of the mood prediction, (2) whether the narrative helped them understand the music, (3) whether the image reflected the music’s emotional character, and (4) which representation they found most helpful overall (label, narrative, image, or all combined). We additionally collect listeners’ music background to explore potential differences across user groups. We acknowledge that this retrospective evaluation does not provide the same level of experimental control as a dedicated comparative study; however, it enables observation of listener responses within a naturally deployed system. At the time of presentation, we will report either preliminary response patterns from the embedded survey or, if data collection is still in progress, the finalized study design and instrumentation details. We hypothesize that generative interpretations may influence listeners’ understanding and engagement with musical emotion differently than categorical labels alone. More broadly, this work investigates how generative AI functions as an interpretive layer between computational music analysis and human perception. Understanding which forms of representation listeners find most meaningful may inform future human-centered music AI systems that adapt explanations and recommendations to different listeners rather than relying solely on fixed emotional categories.