MindTopo: Can Foundation Models Reason in Topological Space?
Abstract
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Piaget and Inhelder’s account places these relations at the foundation of spatial understanding, preceding Euclidean and projective relations, yet foundation-model evaluations focus largely on metric or viewpoint-dependent relations. We introduce MINDTOPO, a developmentally grounded benchmark of topological intuition across five properties adapted from Piaget’s account, subsequent cognitive work, and formal topology: continuity, separation, order, enclosure, and knots. Each property is probed at two cognitive levels. Reasoning asks a model to identify a topological relation or infer how it changes. Planning instantiates a foundation model as a closed-loop agent whose actions must build, preserve, or alter that relation. MINDTOPO contains 11,035 instances across 13 procedurally generated task types with controllable difficulty. Across 13 multimodal large language models, every model scores higher on reasoning than on planning, and the best model remains far below observed human performance. On Qwen3-VL-2BInstruct, supervised fine-tuning and reinforcement learning improve reasoning far more than planning, and agents augmented with image or video generation reach plausible endpoints without preserving topology across transitions.