PolicyAttention: Robustness and Failure Modes of Softmax Policy Improvement
Yuhe Sui ⋅ Yingzhi Tang ⋅ Xinyue Wu
Abstract
Softmax policy updates can be exact in principle yet brittle in causal implementations because finite attention leaks across state groups. We ask when in-context policy mirror descent (PMD)--TD execution remains reliable under routing, critic, and sampling error. First, we give a finite-horizon returned-policy certificate: bounded actor residual $\zeta$ and critic residual $\delta$ yield an explicit suboptimality envelope, with tolerances fixed from step-size and packet certificates before decoder weights are frozen. Second, a matched two-state stress family reveals an exact boundary: implemented improvement is positive iff $\eta c<(1-\gamma)\kappa$, where $\eta$ is the PMD step, $c$ the reward scale, and $\kappa$ the routing margin. At $\kappa=8,\gamma=0.8$, raising $\eta$ from $0.8$ to $2.4$ increases ideal gain from $0.950$ to $2.084$ but flips implemented gain from $+0.916$ to $-2.009$ with an oracle critic. A separate learned pre-LN stress screen shows why averages are insufficient: low-entropy inputs have mean PMD error $0.059$ but sample maximum $1.977$. Thus the same mechanism yields a certified robust regime and a concrete failure boundary, separating nominal policy improvement from reliable implementation.
Chat is not available.
Successful Page Load