Trust What Matters: Language-Conditioned Evidence Routing for Video-IMU Action Question Answering
Tengjun Ni ⋅ Xin Yuan ⋅ Shenghong Li ⋅ Kai Wu ⋅ Wei Ni ⋅ Ren Ping Liu ⋅ Y. J Guo
Abstract
Multimodal action understanding is commonly framed as a fusion problem: combine the available sensors and predict an action label. Synchronized video-IMU question answering (QA) exposes a new and fundamental challenge: The modality that matters more is not fixed. A question about scene context may be answered from appearance, a question about motion dynamics may depend on inertial signals, and a question under occlusion, corruption, or temporal misalignment may require deciding which modality to trust. Thus, the core problem is \emph{language-conditioned evidence selection under unreliable cross-modal observations}, as opposed to multimodal fusion. We propose a structured evidence-to-Large Language Model (LLM) framework for video-IMU action QA. Given synchronized video and Inertial Measurement Unit (IMU), our model first extracts visual, sensor, and fused evidence using modality-specific encoders and bidirectional cross-modal attention. Each candidate answers the queries by independently querying these evidence sources, producing candidate-conditioned visual, sensor, and fused evidence summaries. Then, a reliability-aware router estimates observation reliability and selects how much each evidence source should contribute before projecting compact structured evidence tokens into a frozen LLM for multiple-choice answer prediction. We evaluate on a controlled MMAct-based action-QA protocol designed to isolate perception, sensor-centric reasoning, temporal ordering, reliability reasoning, and cross-modal complementarity. Our method achieves $61.20\%$ accuracy on the held-out test split and $57.00\%$ on the Cross-Modal Challenge subset, outperforming text-only, unimodal, naive fused-to-LLM, and shared-state evidence baselines. Ablations show that structured source identity, candidate identity, reliability tokens, and routing are all essential. Perturbation and shortcut analyses confirm that the model relies on paired multimodal evidence rather than language priors or answer-position bias. These results suggest that robust multimodal action reasoning requires moving beyond fixed sensor fusion toward query-specific, reliability-aware evidence organization for language models.
Chat is not available.
Successful Page Load