UGGRH: Unsupervised Generative Completion and Graph-attention Refinement for Incomplete Cross-modal Hashing
Abstract
The proliferation of multimodal data has underscored the imperative for efficiency in cross-modal retrieval. By mapping data into a unified Hamming space, Cross-modal hashing (CMH) significantly enhances both storage and computational efficiency, thereby establishing itself as a preferred favored paradigm for retrieval tasks. However most methods assume complete data, which rarely holds in practice. In incomplete settings, missing modalities (e.g., text or image) break semantic correspondence and make unsupervised hashing unreliable. Therefore, we propose Unsupervised Generative Completion and Graph-attention Refinement for Incomplete Cross-modal Hashing (UGGRH) method. Specifically, UGGRH leverages pretrained generative models for modality completion and employs a discriminator-guided similarity refinement module to fuse intra-modal relations with paired image-text matching signals, thereby constructing a robust cross-modal similarity matrix that yields a noise-tolerant similarity target for training. After that, it learns discriminative binary codes via hash-aware graph attention and cross-modal fusion, to refine modality embeddings and produce unified hash codes. Experiments on MIRFlickr25K, NUS-WIDE, and MS COCO datasets demonstrate that UGGRH consistently achieves superior performance under high missing rates and modality imbalance, showing significant improvements in retrieval accuracy and robustness across challenging incomplete scenarios.