A minimal recurrent model captures ventral-stream geometry
Abstract
What computations are sufficient to produce ventral-stream-like visual representations? We test a simple hypothesis: repeatedly applying the same learned transformation. We train compact weight tied convolutional networks in which an identical residual block is applied 1, 2, 4, or 8 times while all other model and training choices are held fixed. Ventral-stream alignment in the Natural Scenes Dataset increases monotonically with repeat count, and the eight-repeat model matches or exceeds several widely used ventral-stream models despite having only ˜150K parameters. Repetition also increases object-manifold signal-to-noise ratio and capacity, producing better-separated and more linearly separable representations. These results show that repeated application of a learned computation is sufficient to produce substantial ventral-stream representational alignment, and suggest a simple computational principle for how visual representations can become progressively more structured across a hierarchy.