Text Barely Beats Timing: Predicting YouTube Replay Peaks from Transcripts
Eshan Arora ⋅ Ashna Arora
Abstract
We ask how much text alone, the transcript and title, reveals about which moments of a video viewers will replay, using YouTube's Most Replayed (MR) signal and $13{,}008$ unique videos in three overlapping tiers: five genre corpora, five single-channel corpora, and five held-out channels. We introduce RAPS, a labeling procedure that converts the MR curve into training segments, and compare a fine-tuned RoBERTa with title cross-attention against chance, position, TF-IDF, and BiLSTM baselines, with and without MR-aligned segments at test time. First, a simple position baseline is strong: peak locations are consistent within genres and channels, and position scores come within $0.03$ mAP of the transformer on most corpora. Second, the transformer improves on it by $+0.022$ to $+0.027$ mAP on average on genre corpora, most where position is weakest. Third, training on RAPS segments rather than uniform (equidistant) segments helps only when test segments are also RAPS-aligned (mean $+0.031$, positive on all ten corpora); without MR at test time the effect disappears. On held-out channels, genre-trained and channel-trained models transfer equally well, both beating position on three of five channels. We conclude that, at this granularity, most of the predictable structure in replay behavior is temporal, with text adding a small signal on top.
Chat is not available.
Successful Page Load