How Scale Shapes Compositionality: A Case Study in Meta-Learning
Abstract
Compositional structure is a crucial component of human thought---it enables us to learn rapidly and generalize effectively when faced with novel combinations of familiar elements. An important question for today’s AI researchers is how to build machines that have compositional abilities. On one hand, many recent advances in AI---including in explicitly compositional domains such as coding---have been enabled largely by scaling systems up; perhaps, then, scale pushes models towards compositional solutions. On the other hand, arguments from cognitive science suggest that compositionality emerges only when a system faces resource constraints that make memorization infeasible. Given these competing intuitions, what is the relationship between scale and compositionality? Motivated by this question, we conduct experiments analyzing compositional structure in both model activations and model parameters. We find that the smallest and largest models that we investigate converge to solutions that are non-compositional or only trivially compositional, but models at in-between sizes display robust structural compositionality. These results point to a nuanced relationship between scale and compositionality: For compositionality to emerge, models must be sufficiently large to represent the relevant compositional solution, but if a model is above that threshold, then capacity constraints can indeed promote compositionality.