View-Spectral Reconciliation Learning for Text-based Multispectral Aerial–Ground Person Re-Identification
Abstract
Conventional text-based person re-identification (ReID) approaches mainly rely on visible-spectrum images captured from ground-level viewpoints. However, with the increasing prevalence of UAVs and multispectral imaging systems, pedestrians can be captured from both ground and aerial viewpoints, and different spectral sensors provide complementary cues that reveal diverse pedestrian characteristics. As a result, conventional text-based ReID methods become inadequate for such retrieval scenarios. To address this limitation, we introduce a novel task termed Text-based Multispectral Aerial-Ground Person Re-Identification (TMAG-ReID), which aims to retrieve pedestrian images across aerial and ground viewpoints under multispectral scenarios using textual descriptions. To facilitate this task, we construct the T-MSAG dataset, which consists of multispectral pedestrian images captured from aerial and ground viewpoints, paired with corresponding textual descriptions. Furthermore, we propose a novel View-Spectral Reconciliation Learning framework (VSRL) that leverages textual guidance to interweave discriminative semantics across distinct views and spectral modalities. This fusion yields robust cross-view and cross-modal pedestrian features, thereby enabling high-performance retrieval. Extensive experiments on the T-MSAG dataset demonstrate the effectiveness of VSRL.