Pixels over Symbols: Sensory Realism Improves Behavioral Alignment in Models of Cognition
Abstract
Computational models of cognition have traditionally relied on abstract symbolic inputs that approximate sensory experiences with simplified vector representations. This simplification reflects historical computational constraints and a long-held assumption that the details of sensory realism are incidental to the study of cognitive processes. Recent advances have enabled the development of sensory-realistic models capable of performing diverse cognitive tasks, yet whether sensory realism meaningfully shapes cognition has not been formally established. We directly address this gap through a large-scale empirical comparison of abstract and sensory-realistic neural network models trained on a battery of memory-dependent decision-making tasks. Our results show that sensory-realistic models mirror human behavioral response patterns significantly more accurately than their abstract counterparts, though a substantial gap to human internal behavioral consistency remains even for the best tested models. By analyzing the internal dynamics of these systems, we demonstrate that the recurrent dynamics in abstract models is not inherently divergent from that in natural models. In fact, frozen abstract-trained weights can efficiently solve cognitive tasks from sensory-realistic inputs, however, at the cost of diverging the dynamics from their original configuration. Together, our findings argue that sensory realism is a primary determinant of human-like behavior in artificial systems, suggesting that findings obtained from abstract task models in neuroscience should be more carefully interpreted in light of their potential behavioral limitations.