PolicyAttention: When Softmax Attention Becomes Policy Improvement
Yuhe Sui
Abstract
Softmax normalization is often viewed as a constraint on implementing reinforcement-learning updates with Transformers. For negative-entropy policy mirror descent (PMD), however, it is exactly the policy-improvement geometry: the classical update is $\pi^+(a\mid s)\propto \pi(a\mid s)e^{\eta Q(s,a)}=\operatorname{softmax}(\log\pi+\eta Q)_a$. Starting from this identity, we build a causal actor--environment--critic computation in context: softmax forms the improved policy, fresh interaction supplies on-policy samples, and a later stage performs TD evaluation. For any prescribed finite horizon, we construct a fixed two-block, two-head normalization-free decoder whose finite routing, critic, and sampling errors propagate to a pathwise bound on the suboptimality of the returned policy; with certified near-greedification, the bound becomes geometric decay to an explicit approximation floor. To connect the control theorem to learning, we show that a PMD-aligned KL-proximal objective selects the same actor in the complete centered linear action-equivariant class, with Fisher-weighted effective-coordinate optimization. As a standard-architecture realization check, an ordinary pre-LN Transformer trained without direct PMD actor labels under the aligned objective achieves held-out PMD row-$\ell_1$ error $0.0357$ and TD MAE $0.121$. Together, the construction, last-policy guarantee, and learned realization establish one representation-to-action mechanism: softmax turns policy and value information into test-time policy improvement, while causal TD evaluation closes the loop from representation to control.
Chat is not available.
Successful Page Load