RoSeViT: Role-Separated Vision Transformers for ARC Visual Reasoning
Abstract
ARC tasks can be formulated as visual transformation problems, where a model must infer a high-level transformation rule from a small set of demonstrations and apply it consistently across a grid. While dense Vision Transformers (ViTs) perform well in this setting, they entangle local visual state and global transformation logic within a shared token representation, which can hinder consistent reasoning. In this paper, we introduce RoSeViT, a role-separated Vision Transformer that explicitly decouples high-level instruction inference from visual workspace computation. RoSeViT uses controller tokens to aggregate transformation-level information and workspace tokens to represent the visual grid, with a structured attention mechanism that routes nonlocal interactions through the controller. This design is motivated by a formal analysis showing that enforcing a shared instruction representation can reduce the effective complexity of visual transformation modeling. Under the standard VARC training and evaluation protocol for ARC, RoSeViT consistently outperforms dense ViT baselines. In a controlled backbone replacement setting, it improves ARC-1 accuracy from 54.5 to 56.6, and with recurrent refinement reaches 58.8. When integrated into the full system, RoSeViT-Ensemble improves performance from 60.4 to 62.6 on ARC-1. These results demonstrate that explicitly separating instruction and workspace representations leads to more effective visual reasoning on ARC tasks.