Mixture of Layers: Dynamic Layer Routing for Visual Reasoning
Abstract
Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely only on the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and aggregates selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. Our proposed MoL is a vision encoder-agnostic framework for layer-wise routing that selectively samples visual features useful for fine-grained visual reasoning tasks. Experiments across seven fine-grained visual reasoning benchmarks demonstrate substantial performance improvements, including +18.9% on V* overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to baseline MLLMs, without requiring multi-resolution inputs, simple interleaving of multiple vision encoders, or additional patch tokens. We further analyze receptive field scales and routing behaviors across vision encoder layers to explain why adaptive layer selection improves perception and reasoning. Our results suggest that conditional intermediate-layer representations are a key step toward stronger visual perception and reasoning in MLLMs. Code will be open-sourced.