Position: Linguistic Structure Reveals the Preconditions for Foundation Modeling
Abstract
Foundation models have been especially successful in natural language, where large-scale self-supervised learning operates over representations with extensive reuse and compositional structure. We argue that this success reflects not only scale, architecture, and tokenization, but also structural properties of language that may generalize beyond it. We put forward a falsifiable position: foundation-model behavior requires recurrent, compositional representations. Recurrent units reappear across contexts with stable roles, allowing learning signals to aggregate, while compositional structure supports systematic recombination beyond observed instances. The claim concerns these structural properties rather than linguistic surface form. We propose operational diagnostics for assessing whether a representation is foundation-ready, including cross-context unit reuse, distributional stability, bigram reuse, novel recombination, and lightweight transfer. Synthetic stress tests in spatio-temporal learning, reinforcement learning, and physical-field dynamics show that raw or coordinate-specific representations often fragment experience into context-specific units, while motion, situation, local-observation, and dynamical-regime abstractions expose more reusable structure and improve transfer. Physical-field dynamics provide a useful contrast, since recurring local structure is already partly present in the data. These results suggest that extending foundation modeling beyond language requires assessing whether the chosen representational medium supports recurrence, compositionality, and statistical reuse.