MUST: Stage-Adaptive Stability Control for Test-Time Scaling in Multimodal Reasoning
Abstract
Test-time scaling (TTS) improves reasoning by allocating additional inference compute, but in multimodal large language models (MLLMs), more compute is not always reliable. Longer reasoning may amplify visual misperception, induce answer drift, reinforce majority traps, or turn initially correct predictions into overthought failures. We propose MUST, a training-free framework that recasts multimodal TTS as stage-conditioned stability control. MUST separates inference into reasoning and verification stages and assigns each stage a different reliability criterion. During reasoning, Contrastive Answer-Manifold Stability (CAMS) selects answers whose supporting trajectories form compact and separable latent groups. During verification, Attention-Free Support-Chain Stability (AF-SCS) evaluates whether explicit visual evidence and subsequent reasoning form a stable latent support chain, without relying on token-to-patch attention maps. Experiments on multimodal reasoning benchmarks show that MUST improves the accuracy-compute trade-off and reduces overthinking-induced errors. These results suggest that reliable MLLM test-time scaling should adaptively control when and how extra compute is trusted, rather than simply increasing inference budget.