Comparing Direct GPT-Audio and Feature-Based GPT-4o-Mini Prediction of Human Emotion Ratings from Transformed Infant Vocalizations
Kael Kameoka ⋅ Takashi Yamauchi
Abstract
Audio-capable large language models can process raw audio, but whether they reproduce fine-grained human emotion judgments for non-linguistic vocalizations is unclear. We compared three strategies on 178 matched transformed infant vocalizations (cooing, laughing, crying, screaming, babbling) rated for happiness, sadness, anger, fear, and disgust: direct gpt-audio prediction from raw audio, supervised prediction from engineered acoustic features, and GPT-based in-context learning using the same features and acoustically similar labeled examples. Direct gpt-audio showed little correspondence with human ratings ($r=-0.173$ to $0.016$), whereas supervised feature-based models reached $r=0.883$. K-NN baselines ($k=25,13$) ranged from $r=0.543$ to $0.789$; GPT in-context learning ranged from $r=0.573$ to $0.786$, closely matching $k=13$ and adding little beyond local acoustic similarity. These findings suggest explicit acoustic representations provide useful task-relevant structure and reveal a gap between general raw-audio processing and fine-grained affective prediction.
Chat is not available.
Successful Page Load