One-Way Coupling via Patchwise Adaptive Normalization in Transformer Surrogates
Abstract
Many physical systems exhibit \emph{one-way coupling}: exogenous fields such as material density or obstacle geometry influence the dynamics without being affected by them. Current transformer-based PDE surrogates either ignore this structure by concatenating the conditioning field to the input channels or reduce it to a single global vector via standard Adaptive LayerNorm (AdaLN), discarding spatial information in both cases. We introduce \textbf{Patchwise Adaptive Normalization} (\texttt{PAdaNorm}), which partitions the conditioning field into patches aligned with the input tokenization and produces spatially varying scale, shift, and gate parameters via a shared MLP encoder. Each transformer block thus receives per-patch modulation that enforces the spatial locality of the one-way coupling at negligible parameter overhead. We evaluate \texttt{PAdaNorm} on acoustic scattering datasets and on 2D Navier--Stokes simulations, using a Walrus-inspired transformer architecture. Against parameter-matched baselines, \texttt{PAdaNorm} reduces long-horizon variance-scaled RMSE by up to 35\% on 60-step autoregressive rollouts, with the advantage amplifying at longer horizons. An ablation over eight conditioning variants shows that modulating a single attention axis suffices and that a lightweight shared MLP remains competitive with per-block projections.