AffordVLA: Implicit Affordance Representation Alignment for Vision-Language-Action Models
Abstract
Recent advances in vision-language-action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, their visual representations can remain dominated by global object appearance and struggle to focus on task-relevant functional interaction regions, limiting robustness in unstructured environments. Explicit affordance inputs can guide manipulation but may require additional annotations and introduce perception dependencies and inference overhead. To address these limitations, we propose AffordVLA, an affordance-enhanced VLA framework that internalizes manipulation-centric affordance perception into VLA visual representations through implicit representation alignment. A frozen zero-shot affordance teacher extracts task-conditioned visual representations from RGB observations and language instructions. During training, these representations supervise intermediate VLA visual features without additional task-specific affordance annotations; the teacher is removed at inference. Extensive simulation and real-world experiments show that AffordVLA achieves average success rates of 61.2\% and 28.8\% on five RoboTwin2.0 tasks in the Easy and Hard settings, respectively, and 87.5\% across eight real-world tasks. Ablations indicate improved manipulation success and faster progress in optimization steps, with comparable policy inference efficiency.