mmLIP: mmWave Radar-Language Interactive Pretraining via Point Confidence
Abstract
Millimeter-wave (mmWave) radar provides core sensing capabilities to traditional vision, particularly under occlusion, low-light, and privacy-constrained conditions. Recent efforts have explored integrating radar signals with large language models (LLMs) for high-level semantic reasoning over the received radar data. However, existing approaches typically project radar inputs into borrowed vision or LiDAR embedding spaces, which are not tailored to radar-specific physical cues such as Doppler and reflection intensity. This design introduces a representational bottleneck, hindering the extraction of informative radar semantics. To overcome this limitation, we propose \textbf{mmLIP}, a novel radar-language interactive pretraining framework that learns radar-specific alignment without relying on external embedding spaces. Our approach directly aligns radar representations with the text embedding space via a point-level contrastive objective, enabling fine-grained correspondence between radar points and textual tokens. In addition, we introduce a confidence-aware contrastive learning mechanism that adaptively reweights radar tokens based on their semantic relevance, promoting informative signals while suppressing clutter. Supported by our newly curated radar-text pairs, mmLIP captures structured radar semantics. Extensive experiments demonstrate that mmLIP successfully integrates with diverse LLMs or vision-language models, achieving state-of-the-art performance on zero-shot bidirectional retrieval as well as text generation tasks, including captioning and question answering.