ReactionBench: A Human-Anchored Benchmark for Modeling Reactions to Video
Yibo Zhao ⋅ Ao Qu ⋅ Xuan Jiang ⋅ Jiahao Xu ⋅ Keane Ong ⋅ Hang Jiang ⋅ Zhaofeng Wu ⋅ Dingyi Zhuang ⋅ Yihong Tang ⋅ Kaichen Zhou ⋅ Jinhua Zhao ⋅ Paul Liang
Abstract
Socially responsive multimodal agents must interpret not only what happens in a video, but also how people respond to it, and understanding a response requires linking it to the event that triggered it. In this work, we consider this problem of stimulus-reaction understanding. We introduce $\textbf{ReactionBench}$, a human-anchored benchmark that measures whether models can rank reaction intensity within a viewing episode, match a reaction to its triggering moment, and explain the stimulus cues behind it. ReactionBench starts from 433 hours of synchronized in-the-wild stimulus-reaction video; an automated, model-assisted selection pipeline narrows this collection to 3,257 candidate moments, each rated at least three times by a pool of eight annotators. Across seven multimodal models, the strongest joint intensity-ranking result remains far below the pool-specific human reference (0.176 versus 0.649), while the best four-way trigger-matching accuracy reaches 50.9%. A paired scoring-protocol comparison shows that evaluating moments independently improves intensity ranking for every model. Supervised adaptation of one open-weight model at this scale does not close the gap: some apparent gains arise from output-format compliance or collapsed predictions rather than improved reaction understanding.
Chat is not available.
Successful Page Load