Propagation of Chaos in Contextual Flow Maps
Chen ⋅ Zhengjiang Lin ⋅ Kaizhao Liu ⋅ Philippe Rigollet
Abstract
We develop a quantitative statistical theory of transformers in the large-context regime by adopting the abstraction of contextual flow maps: dynamical systems that evolve a distinguished token in the presence of a contextual measure across a stack of attention blocks. Within this framework, the finite-context model approximates an idealized infinite-context system in which the contextual measure is replaced by its underlying population, so that the context length $n$ becomes a statistical resource. Exploiting the McKean--Vlasov structure of the dynamics and the classical machinery of propagation of chaos, we establish a forward bound controlling the deviation between the finite- and infinite-context flow maps uniformly along depth, and a backward bound controlling the deviation between the corresponding training trajectories uniformly across iterations of online gradient descent. Both bounds achieve the optimal Wasserstein rate $n^{-1/d}$. The analysis rests on a new Eulerian adjoint formulation of the loss gradient and stability estimates for the resulting forward--adjoint system, both of which may be of independent interest. We further verify that the standard transformer architecture satisfies these estimates with Lipschitz constants independent of embedding and parameter dimensions, suggesting a structural reason why transformers train stably across model scales.
Chat is not available.
Successful Page Load