Why Take Every Step? Temporal Striding in Autoregressive PDE Surrogates
Abstract
Autoregressive surrogate models are usually trained to predict one stored simulation timestep ahead. We study temporal stride (s), the number of stored timesteps between successive model states, as a training and inference choice rather than a fixed sampling convention. Across two datasets from The Well (PlanetSWE and Shear Flow) and three architectures (TFNO, AViT, and ConvNext U-Net), we isolate the effects of training stride, evaluation stride, and explicit stride conditioning. We find that conventional stride-1 setup is not the best-performing configuration for any architecture-dataset pair. Models trained on strides 1-4 also generalize surprisingly well beyond this range: unseen strides 5-8 remain stable in most cases and often reduce long-horizon rollout error, with stride 8 frequently performing best. Conditioning avoids TFNO divergence on PlanetSWE and improves AViT at all strides tested. For AViT, larger evaluation strides also suppress patch-frequency artefacts in rollout residuals. These results identify temporal stride as an underanalyzed but consequential design choice for autoregressive surrogate models. As autoregressive PDE models grow in scale and scope, we hope these findings offer useful guidance for how temporal stride is chosen during training and inference.