Rotate the Loop: Restricted Injections for Content Retention
Alexey Zemtsov ⋅ Alexander Derevyagin ⋅ Albina Klepach ⋅ Vadim Lomshakov ⋅ Andrei Polubarov ⋅ Nikita Lyubaykin ⋅ Vladislav Kurenkov ⋅ Alexander Nikulin
Abstract
Looped transformers reuse one shared block across many passes, so depth becomes a quantity that can be spent at runtime rather than a stack of distinct layers. Theory and practice both push this iteration toward a fixed point, yet a loop that contracts too hard can forget what a later pass still needs. We show that retention and stability are one geometric question: the fraction of a write that survives in the worst direction, relative to the best, is exactly the condition number to the power of the budget, independently of the radius. What is relevant is the relative floor, available only if the condition number itself is boundable at design time. The maps that make it so are a neighbourhood of the conformal orthogonal group: one shared scale is the perfectly even case, and a few rates remain a ratio of a few numbers, whereas a free diagonal has no such bound. A pairwise rotation on the loop axis implements that family at no material cost in memory or compute, and without disturbing the rest of the model. The same change helps on $A_5$ word problems, where the product must survive until the last token, and in preliminary language-modeling experiments.
Chat is not available.
Successful Page Load