MLLM Makes Strong Backbone for Multi-Modal Object Detection
Abstract
Multimodal large language models (MLLMs) provide a promising alternative by taking object detection as a sequence generation task. However, applying MLLMs to multimodal object detection poses challenges in visual cross-modal (RGB-Infrared) fusion and coordinate prediction under Cross-Entropy supervision. In this paper, we propose a novel MLLM-based object detection framework that addresses these challenges through improving visual cross-modal token alignment and coordinate generation. Specifically, we formulate RGB-IR fusion as pre-decoding token-space modality routing, implemented by the Cross-Modal Complementary Adapters and the Visual Token-wise Router. We further reformulate coordinate-token learning as scale-adaptive neighborhood likelihood maximization through our proposed Geometry-Aware Loss, bridging discrete token prediction and continuous geometric localization. Extensive experiments show that our method achieves state-of-the-art performance, reaching 67.81\% mAP@0.5 under in-distribution scenarios, outperforming existing baselines by over 8\%, while also exhibiting strong generalization to out-of-distribution and referring object detection tasks.