Sparse Planning in Visual World Models via Cost Gradients
Yingchen Xu ⋅ Edward Grefenstette
Abstract
Token-based world models enable fine-grained latent planning, but iterative search inside the predictor is dominated by the number of spatial tokens processed during rollout. We introduce COSTGRAD, a training-free, goal-conditioned selector that ranks spatial tokens by the gradient norm of the planning cost with respect to each input token. By deriving importance from the downstream control objective, COSTGRAD targets tokens that matter for planning rather than merely for prediction. On AdaLN-conditioned predictors at $50\%$ sparsity, COSTGRAD matches or exceeds full-token planning on three of four continuous-control benchmarks, while giving a measured $2.6\times$ wall-clock speedup per environment planning step. The sparsity dividend can be combined with reduced CEM search, yielding a $\sim5\times$ total speedup while still exceeding the full-token baseline. We also identify an architecture-dependent failure mode: in a matched AdaLN-vs-concat comparison, concat maintains comparable full-token performance but pure COSTGRAD loses its advantage over random selection. This failure tracks action-pathway drift under selected token removal: AdaLN largely preserves how actions affect kept tokens, whereas concat perturbs the action-conditioned computation. Mixing in random anchors partially recovers COSTGRAD's advantage on some matched-concat checkpoints, suggesting selector--architecture compatibility as a design axis for sparse world-model planning.
Chat is not available.
Successful Page Load