Reranking with Intra-modal Visual Association for Text-to-Image Person Re-Identification
Abstract
Text-to-image person re-identification (TIReID) aims to retrieve images of one pedestrian using natural-language descriptions. It remains fundamentally challenging due to significant inter-modal semantic gap between sparse textual cues and fine-grained visual appearance. Existing methods mainly perform instance-wise text-image matching during inference, overlooking intra-modal visual associations among top-ranked gallery candidates, where multiple images may depict the same identity. To exploit such associations, we propose Reranking with Intra-modal Visual Association (RIVA), a plug-and-play reranking framework that leverages intra-modal visual association among retrieved candidates to guide multimodal large language model (MLLM)-based reranking. RIVA consists of two complementary components. First, we propose an MLLM-based Short-lIst Visual Association (SIVA) method, which is trained with a three-stage, dual-task curriculum to learn identity-level image-image association and then use it as visual context for text-image matching. Second, we develop a long-list visual association approach based on clustering, which clusters the top-ranked candidates in the feature space of pedestrian images, uses SIVA to examine the most likely matched cluster, and calibrates the image-text similarity scores accordingly. Extensive experiments on three popular TIReID benchmarks demonstrate that RIVA consistently improves retrieval accuracy in multiple evaluation settings. The code will be released.