OmniSE: Support-Faithful Optimization for Evidence-Centric Audio-Video Reasoning
Abstract
Omni-modal models are expected to reason over synchronized visual events and audio cues, yet answer accuracy alone can conceal failures in where an answer is grounded and which modality supports it. In audio-video question answering (AVQA), models may produce correct answers while relying on incorrect temporal spans, dominant-modality shortcuts, or unsupported cross-modal relations. We argue that this failure arises from treating reasoning as a completion-level prediction problem, where supervision and reward signals do not distinguish between support selection and reasoning over the support. To address this, we introduce OmniSE, a support-faithful optimization framework that makes question-conditioned audio-video support the central training object. Each QA instance is rewritten into an Audio-Visual Evidence Trace, a compact optimization unit that records temporal evidence, modality-specific cues, cross-modal relations, and cited reasoning steps. Given this support object, OmniSE first uses Support-Tree Rollout (STR) to explore alternative support hypotheses before sampling support-conditioned reasoning and answers. It then applies factorized reward assignment to score temporal localization, modality evidence, cross-modal binding, reasoning faithfulness, and answer correctness at their corresponding comparison levels. Finally, reward-gated self-consistency regularizes sibling continuations under reliable support, encouraging fixed trustworthy evidence to induce stable reasoning and answers. Experiments on multiple AVQA benchmarks show that OmniSE consistently outperforms strong open-source baselines at the same scale, demonstrating the importance of support-centric optimization for evidence-centric audio-video reasoning. Code and data would be publicly available after peer review.