One Loss Is Not Enough: Cross-Objective Structured Pruning for Reasoning Vision–Language–Action Models
Abstract
Reasoning vision--language--action (VLA) models couple language-based reasoning with continuous action generation, but this dual-output structure creates a challenge for model compression: which objective should determine what is safe to prune? We show that neither one alone is sufficient. In Alpamayo-style driving VLAs, cross-entropy (CE) supervises textual reasoning while flow matching (FM) supervises continuous trajectory generation, and the two objectives expose complementary structural importance in the shared VLM backbone. We introduce \textbf{Cross-Objective Structured Pruning (COSP)}, a one-shot method that estimates the importance of attention heads and MLP channels under both objectives and preserves a structure whenever either reasoning or action identifies it as important. Under an exactly matched 24.0% parameter reduction on Alpamayo-1.5-10B, single-objective pruning exhibits complementary failures: FM-only pruning retains competitive driving performance but substantially degrades reasoning, while CE-only pruning severely degrades trajectory prediction and closed-loop driving. COSP achieves the strongest joint performance across PhysicalAI-AV open-loop trajectory prediction, LingoQA reasoning, and AlpaSim closed-loop driving among the evaluated pruning methods. It retains 68.8 LingoQA while improving the dense model's closed-loop scene score from 0.750 to 0.828 and reducing the at-fault collision rate from 0.040 to 0.023. These results show that effective compression of reasoning VLAs requires structural importance to account for both reasoning and action objectives.