To Think or Not to Think: Pre-Decisional Reasoning Budgets for Referring Audio-Visual Segmentation
Abstract
Multimodal reasoning systems increasingly assume that longer chain-of-thought uniformly improves downstream grounding. We show that this assumption fails in referring audio-visual segmentation: simple queries can be harmed by excessive reasoning, creating an “overthinking trap” that dilutes visual focus, while genuinely ambiguous queries still require multi-step reasoning. To systematically study this, we construct counterfactual reasoning-budget labels under a rigorous, leakage-free evaluation protocol. Our results recast adaptive multimodal reasoning as a pre-decisional representation problem: before the first generated token, the generation-onset state contains linearly decodable information predictive of whether additional reasoning will improve segmentation. A linear probe predicts binary reasoning need with 70.1% accuracy, substantially outperforming text-only features and uncertainty-based proxy signals. As a router, the onset-state controller improves over downstream confidence routing and closely matches Short-then-Decide while using far fewer tokens. Mechanistic decomposition suggests that this signal is better captured by distributed high-dimensional representations than by isolated manual heuristics. Reusing this onset state for lightweight, single forward-pass routing reduces autoregressive token consumption by 60% while retaining ~96% of always-long performance. It further transfers to out-of-distribution (OOD) benchmarks without target-split tuning. Comprehensive error analysis shows that bypassing unnecessary reasoning can reduce phrase and localization drift. Ultimately, our results indicate that adaptive multimodal reasoning can be reliably and efficiently routed from pre-decisional internal representations.