Towards Generalizable Partially Relevant Video Retrieval
Abstract
Partially Relevant Video Retrieval (PRVR) aims to retrieve an untrimmed video that contains at least one segment relevant to a text query. Despite its practical motivation, existing PRVR methods are typically trained and evaluated on a single curated benchmark, limiting insight into whether they learn generalizable fine-grained text-video alignment. To examine this issue, we introduce a multi-source-to-unseen-target protocol, in which multiple source benchmarks are merged into a single training corpus without source labels and the trained model is evaluated on a held-out target benchmark. Under this protocol, existing PRVR methods generalize poorly to unseen targets despite leveraging the general image-text alignment of CLIP. Surprisingly, they even underperform zero-shot CLIP, suggesting that PRVR training can actively erode the transferable alignment inherited from CLIP. To address this limitation, we propose Generalization-Aware Learning (GAL) for PRVR, a two-branch framework that preserves CLIP-derived transferability while learning PRVR-specific fine-grained text-segment matching. GAL consists of an anchored branch that maintains CLIP's transferability and an adaptive branch that learns task-oriented segment-level alignment. We train these branches through conservative residual updates, per-sample branch specialization, and CLIP-text relation distillation, enabling complementary learning without overwriting the alignment structure provided by CLIP. Under the proposed protocol, GAL achieves state-of-the-art unseen-target retrieval with competitive source performance, showing that fine-grained matching can be learned without collapsing transferable alignment.