Spectral Latent World Models for Cadence-Flexible Planning
Abstract
Latent world models for planning are grid-locked in time, and removing the sequential rollout does not fix it: an autoregressive predictor trained at frameskip 5 emits latents only on its 5-frame grid, and predicting all action-prefix latents in one parallel pass instead lands them on the same five block boundaries. A controller replanning on a finer clock has nothing to score between hops, and must round or interpolate. We present Spectral WM, which emits a whole chunk of future latents in one forward pass as temporal DCT coefficients in a basis fixed to the window's full raw-frame resolution, a curve rather than a sequence of hops, so a state at any raw-frame offset is a single inverse-transform readout at identical cost. Temporal resolution becomes a property of the readout, not the architecture. One frozen model holds its success as replanning is refined from every 25 frames to every frame, losing at most 2.9 pp on any environment, leading the strongest autoregressive off-grid strategy at every-frame replanning by +6.1 pp on average and up to +10.2 pp. Grid-locked predictors reach parity only through an interpolation readout we supply for them, which the spectral curve does not need. The capability is free elsewhere: as a drop-in swap on a frozen LeWM encoder, Spectral WM leads Fast-LeWM (91.2 vs. 90.5) on average success and plans ~3.6× faster than LeWM. We also isolate a predictor-independent prerequisite for fine-cadence replanning: the scoring deadline must be stationary, or success falls by 7–65 pp across tasks regardless of the predictor.