Three formal language tiers of artificial and biological intelligence
Abstract
Although computational abilities in AI can arise in part from scaling data and model size, the underlying machinery necessary for language and reasoning capabilities remains poorly understood. Via principled analyses we show evidence for differing fundamental abilities among specific classes of language model architectures, even with extensive training. Specifically, new computational abilities arise from transitions between computational tiers in the grammar-automata hierarchy from the Theory of Computation. We provide empirical evaluations of both language models and human subjects, culminating in a theoretically-grounded, pragmatic benchmarking method that can be used to ascertain the computational tier of a computational system: artificial or biological. Tiers for the transition from prelinguistic to fluent natural language are identified, and the tiers for the subsequent transition from articulate language use to reliable logical reasoning are discussed. The results offer an explanatory account of the abilities and shortfalls of a range of large language models, suggesting actionable insights into the expansion of their logical reasoning capabilities.