TF-PRVR: Training-Free Partially Relevant Video Retrieval for Real-World Generalization
Abstract
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and weak cross-domain robustness due to task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR designed to improve real-world generalization. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments on standard PRVR benchmarks demonstrate competitive retrieval performance and strong robustness under cross-domain evaluation, suggesting a practical direction for real-world PRVR.