MatchSense: Multimodal Reasoning for Key Moment Localization in Long-Form Sports Broadcasts
Abstract
Long-form sports broadcasts often contain only a few decisive events hidden within long periods of regular gameplay, making key moment identification challenging for existing video-language models. To address this, we introduce MatchSense, an audio-guided language reasoning framework designed for identifying key moments in soccer broadcasts. MatchSense first leverages prosodic saliency to detect Salient Object Intervals (SOIs), which significantly reduce the temporal search space before applying structured video-language reasoning using and outputs. To effectively handle varying numbers of events and diverse reward signals, we adapt Group reward-Decoupled Normalization Policy Optimization (GDPO) for assignment-based key moment prediction. In this process, predicted events and ground-truth events are aligned using Hungarian assignment, after which multiple reward components including temporal overlap, semantic alignment, output-format validity, and temporally aligned acoustic saliency are normalized independently before policy optimization. Under the proposed video-language key moment identification setting, MatchSense achieves an mAP of 9.42%, reduces the mean boundary error to 1.54 seconds, and improves Precision@0.5 from 6.80% to 11.50% compared to SFT-only training. Additionally, the audio-guided pruning strategy retains only 35.3% of the original timeline while preserving 83% of salient events, resulting in a 64.7% reduction in FLOPs and a 2.8× inference speedup. The code and implementation details are publicly available through an anonymous repository.