ClutterBench: A Benchmark for Evaluating Vision-Language-Action Models in Cluttered Environments
Abstract
Humans naturally perform manipulation in cluttered environments, reasoning about occlusions, collisions, and constrained motion. However, existing robotic manipulation benchmarks and learning pipelines predominantly evaluate performance in sparse and structured scenes. These settings now serve as the primary training and evaluation source for modern Vision-Language-Action (VLA) models. As a result, the behavior of VLA models in cluttered environments remains poorly understood. We introduce ClutterBench, the first benchmark designed to assess the manipulation capabilities of VLA models in cluttered environments. ClutterBench provides four task categories: orderly, random, complex, and temporal dynamic clutter. We evaluate 7 mainstream state-of-the-art VLA models in simulation and 3 in the real world, and find that even the best model achieves below 60% average success rate across all clutter types, revealing fundamental limitations in task understanding, spatial reasoning, and interaction planning. We also evaluate trajectories across multiple metrics, including non-target object displacement, enabling finer-grained assessment of behavior quality beyond success rate. Real-world evaluation confirms that model ranking and dominant failure modes are consistent across platforms. To support further research and adoption, we will open-source our benchmark, evaluation code, and datasets upon acceptance.