One Loss Is Not Enough: Cross-Objective Structured Pruning for Reasoning Vision–Language–Action Models
Abstract
Reasoning vision--language--action (VLA) models couple textual reasoning with continuous action generation, making model compression inherently multi-objective. We show that cross-entropy (CE) and flow matching (FM) expose complementary structural importance in the shared VLM backbone, so neither objective alone is sufficient for structured pruning. We introduce \textbf{Cross-Objective Structured Pruning (COSP)}, a method that scores attention heads and MLP channels under both objectives and prioritizes structures important to either. At a matched 24.0\% parameter reduction on Alpamayo-1.5-10B, COSP achieves the strongest joint performance across PhysicalAI-AV, LingoQA, and AlpaSim among the evaluated pruning methods, retaining 68.8 LingoQA while improving closed-loop scene score from 0.750 to 0.828 and reducing at-fault collision rate from 0.040 to 0.023.