VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
Abstract
Despite recent successes in visual reasoning, it is still poorly understood how vision-language models (VLMs) integrate perception, memory, and actions in visual interaction, a key capability that underlies real-world applications such as agents and robotics. Towards understanding these behaviors, we introduce VisGym, a gymnasium of 17 environments spanning symbolic puzzles, image understanding, navigation, and manipulation for the diagnostic evaluation of VLMs across domains. Beyond benchmarking for task performance, VisGym provides flexible controls on difficulty, input representation, planning horizon, and feedback, which enable controlled analyses of how interactive design choices affect model performance. We perform several experiments showing that even the strongest frontier models struggle in visual interaction, achieving low success rates in both the easy (26.8%) and hard (12.6%) configurations. Specifically, our analyses reveal four key vulnerabilities in current models: (1) they fail to effectively leverage long context, performing worse with unbounded history; (2) they struggle to utilize pure visual feedback, degrading when textual feedback is removed; (3) several text-based symbolic tasks become substantially harder once rendered visually; and (4) incorporating explicit goal observations can paradoxically backfire. Finally, each environment includes a heuristic multi-step solver that enables post-training case studies on supervised fine-tuning, including analyses of in-domain improvements, module-level contributions, data curation strategies, and preliminary studies of downstream transfer. We open-source all environments, data generators, evaluation pipeline, and training code at https://anonymous.4open.science/r/VisGym