Semantic Consistency of Vision Tokens: A Vision-Centric Perspective on Multimodal Large Language Models
Abstract
In the computer vision world, vision-language and self-supervised encoders represent two unique paradigms characterized by global language alignment and high-quality vision-only patch representations, respectively. Multimodal large language models (MLLMs) often use the former because of language alignment, despite the recent rise in vision-centric tasks resembling traditional dense vision tasks that would likely benefit from the high-quality patch representations of the latter. This raises the question whether self-supervised models could aid MLLMs in vision-centric tasks. To investigate this, we define the concept of semantic (in)consistency of patch tokens produced by language-aligned encoders and identify two sets of tokens, semantically consistent (SCTs) and semantically inconsistent (SITs). Through systematic probing via the removal of SCTs or SITs from the input of modern MLLMs, we show higher reliance on SCTs than SITs, as removal of the former induces larger performance drops. This demonstrates that MLLMs benefit from semantic consistency of vision tokens. With this in mind, we devise a simple yet effective extension of the MLLM training objective that combines traditional language modeling with semantic guidance via minimization of semantic inconsistency of vision tokens. This strategy, named Semantically-Guided Training (SGT), combines the signals of both loss terms to push the vision encoder towards providing semantically consistent tokens, while preserving language alignment. Through extensive evaluation, we show that SGT drastically outperforms all baselines and MLLMs with specialized vision encoders on all vision-centric benchmarks while keeping strong scores on text-centric benchmarks, thus setting state of the art when averaged across all tasks. Our results highlight the importance of combining language alignment and semantic consistency through the lens of MLLMs rather than independently from them and pave the way towards the design of better MLLM vision encoders.