Modality-Depth Routing for Visual Reasoning in VLM Post-Training
Abstract
The prevailing route to stronger visual reasoning in vision-language models (VLMs) builds costly Chain-of-Thought corpora or curated step-by-step rationales, concentrating progress in well-resourced industry labs. We ask whether the reasoning signal already present in generic visual instruction data can instead be unlocked through a structural change to the post-training interface itself. At fixed data and capacity, standard supervised fine-tuning (SFT) on a generic instruction mixture partially improves reasoning but homogenizes the modality-depth pathways through which reasoning and perception inputs flow. We isolate two coupled patterns. First, visual and textual representations couple at the deepest layer: cross-modal Centered Kernel Alignment (CKA) rises by roughly 12% over the frozen baseline. Second, perception and reasoning inputs are routed through similar late-depth profiles, leaving no mechanism for per-input depth allocation. We propose Modality-Depth Attention (MDA), a lightweight side pathway that adds modality-specific query projections and a sigmoid-gated learnable depth residual. Under a matched protocol on Qwen3-VL-2B, MDA improves the reasoning average over SFT by +3.9 and keeps perception almost intact, at 0.05% extra parameters. Multiple mechanistic probes converge on the same routing account. The gains extend to Qwen3-VL-8B and transfer to LLaVA-OneVision-7B, suggesting that modality-depth routing may be a general lever for unlocking reasoning from generic instruction data without compromising perception. Code will be released upon publication.