PolyVision: Conditional Visual Scaling Via Dynamic Expert Routing For Vision-Centric MLLMs
Tiehan Fan ⋅ Chen Zhao ⋅ Nikai Du ⋅ Zili Yi ⋅ Jian Yang ⋅ Ying Tai
Abstract
Multimodal large language models have achieved remarkable progress, yet their language-dominant designs constrain visual perception and reasoning. Vision-centric multimodal large language models aim to restore visual understanding as the structural foundation, but existing systems rely on static visual expert fusion, leading to redundant computation, weak adaptivity, and limited multimodal transferability. Prior VC-MLLMs scale visual capacity by adding experts, whereas $\mathtt{PolyVision}$ scales visual capacity by allocating experts. We introduce $\mathtt{PolyVision}$, a vision-centric framework for conditional visual scaling in which dynamic visual expert routing assigns image patches to heterogeneous visual encoders. An \textbf{Attentive Router} selectively activates relevant experts for each patch, while a compact \textbf{VisPack} module executes routed representations efficiently and preserves example independence. Together with additional visual pre-alignment, this design turns heterogeneous visual scaling into a practical conditional-computation process. Built on the \texttt{Qwen3} backbone, $\mathtt{PolyVision}$ achieves superior efficiency and scalability on multimodal benchmarks. Experiments show that adaptive expert routing strengthens visual specialization, clarifies practical properties of visual scaling, and improves visual understanding at a favorable trade-off between efficiency and performance.
Chat is not available.
Successful Page Load