RILA: A Radar-Native Structured Interface from Sparse mmWave Point Clouds to Large Language Models
Abstract
Millimeter-wave radar is promising for privacy-preserving human sensing, yet current radar-to-language pipelines typically map sparse point clouds into either global clip features or discrete token codes that are only weakly matched to the structure expected by large language models (LLMs). We present RILA, a radar-native structured interface layer for large-language-model adaptation, which converts sparse mmWave radar point clouds into LLM-readable event-aware tokens. RILA combines kinematic phase-space tokenization, flow-aware dual-order serialization, dual-scan selective state-space encoding, and an event-aware token abstraction that exposes both global clip context and temporally localized motion units to the language model. The same interface supports three language tasks through a shared decoder: clip summary generation, ordered event description, and temporal question answering. To stabilize the interface before instruction tuning, we introduce a structured interface-language alignment objective over clip and event tokens. We evaluate RILA using physics-aware synthetic pretraining and real-world mmWave adaptation under controlled radar-native baselines and matched decoder settings. This formulation reframes radar-to-LLM modeling as an interface design problem and studies how radar-native structure, rather than only generic cross-modal projection, determines whether sparse mmWave point clouds can be effectively understood by LLMs.