One-shot Familiarity Learning is Facilitated by Visual Category Structure Learned through Training Augmentation
Abstract
The medial temporal lobe (MTL) is critical for one-shot learning of episodic memories. Examining the human MTL at single neuron resolution reveals a high proportion of neurons selective for the visual category of to-be-encoded stimuli, which together with orthogonal memory selective neurons establish a geometric representation of memory. We asked whether one-shot learning is facilitated by category representations. We used a feedforward neural network with anti-Hebbian synaptic plasticity that is capable of one-shot learning and asked whether category representations emerge as a consequence of the need for rapid learning. We trained the network on natural image embeddings on a familiarity recognition task. Training with image augmentations improved familiarity recognition performance at longer delays and produced hidden unit response properties that more closely matched MTL recordings. In particular, such models included both novelty and familiarity selective memory selective neurons, whereas models trained without augmentations did not. Across different trained networks, performance of a given network was correlated with the ability of the model representation to predict visual categories, an attribute that was not part of the loss function used for training. Together, these results suggest that a requirement for one-shot familiarity recognition shapes the representational geometry of intermediate representations, resulting in transformed views of episodes into structured manifolds who take advantage of semantic category structure. These results show a potential benefit for the coexistence of visually selective and memory-selective neurons in the MTL and supports the idea that visual representations are geometrically shaped by the computational demands of one-shot learning. Code: https://anonymous.4open.science/r/emergent-category-EE3C/