Finding Without Words: Re-Identification from Incomplete Multimodal Evidence
Abstract
Re-identification systems typically assume consistent modalities, complete records, or high-quality visual evidence at both query and retrieval time. These assumptions break down in real-world settings where evidence is fragmented, heterogeneous, and evolves over time. Inspired by findings from animal cognition suggesting that perception and identity recognition can occur without symbolic language, we introduce Cognition-Grounded Multimodal Re-Identification (CGMR), a framework for identity retrieval from incomplete visual, acoustic, contextual, and temporal evidence. CGMR combines uncertainty-aware modality encoders, availability masks, and reliability-gated fusion to reason explicitly about missing information rather than treating missing modalities as absent observations. By modeling modality availability, signal reliability, and temporal context, CGMR reframes re-identification as a missing-modality reasoning problem rather than conventional image matching. We formulate a set of falsifiable hypotheses regarding robustness under modality mismatch, visual ambiguity, and temporal drift, and outline an experimental agenda for evaluating multimodal identity retrieval in realistic environments. More broadly, this work argues for identification systems that preserve identity across heterogeneous signals while remaining locally adaptable, privacy-conscious, and human-centered.