What Breaks Depth Extrapolation in Looped Language Models?
Abstract
Looped models offer an alternative path to test-time scaling through weight-tied recurrence, yet their ability to depth extrapolate, which involves solving harder problems by running for more iterations during inference, has largely been demonstrated only on synthetic tasks. Looped language models trained at scale currently struggle with this, and in this work, we study whether the task structure and training objective of language modeling hinder depth extrapolation. First, we introduce a task ladder that gradually moves from pure symbolic composition towards natural language by introducing complex grammatical form and paraphrasing through layout diversity. Grammatical form and length still allow for depth extrapolation, while layout diversity imposes a large convergence cost on learning, and per-example layouts prevent learning the task altogether. Next, we find that the standard next-token prediction training objective for language model pretraining creates a major optimization barrier to depth extrapolation under fixed budgets, while reweighting even 25% of the loss on answer tokens restores it. Our results suggest that future training of looped language models should control layout diversity in the pretraining data and focus the loss on task-important tokens.