Why Transformer-Based Language Models Need Explicit Mechanisms of Cognitive Control
Abstract
We argue that cognitive control, the capacity to hold multiple divergent candidate responses and reconcile them through categorical inhibition rather than soft blending, should be a first-class computational object in transformer-based language models. We develop this claim at two levels. At the \emph{signal level}, interference paradigms reveal a characteristic failure signature: rather than adjudicating between competing responses, current transformer-based language models blend them, and accuracy collapses as interference scales. At the \emph{goal level}, sustained agency requires maintaining concurrent goals, enforcing priority during signal level adjudication, and revising the goal set as evidence accumulates about whether those goals are being successfully pursued; current systems approximate this only through external scaffolding. The architectural gap revealed at the signal level is the same gap that will block transformers from scaling to the kind of goal modulated and self regulating cognition that long-horizon autonomous agents require. We define concrete criteria for what counts as architecturally explicit cognitive control, show that neither the dominant training pipeline nor reasoning-token training implements it, and commit to predictions that distinguish our position from scale-based optimism.