LePlanner: An Iterative Amortized Controller For World Models
Saksham Bansal ⋅ Chayan aggarwal ⋅ Om Naphade ⋅ Vrishin M
Abstract
World models trained with joint-embedding predictive architectures learn compact, structured latent manifolds from physical interaction, yet planning in those latents still relies on one of two costly families of methods. Search-based planners (CEM, MPPI, iCEM) optimize action sequences through many predictor rollouts, achieving strong success at the price of high per-decision compute and latency. Policy-based methods (behavior cloning, GC-IDM) amortize inference into a single forward pass, but degrade on contact-rich tasks where the demonstration distribution is multi-modal. Neither family offers both low latency and robust goal-reaching in a learned latent space. We propose LePlanner, an amortized iterative controller that refines latent action sequences to reach a goal state, trained with an arrival--hold loss that rewards reaching the goal at the earliest feasible horizon and staying there, plus an action gaussian loss which keeps actions in their true manifold. Across goal-directed environments spanning navigation, contact-rich manipulation, and continuous control, LePlanner matches or exceeds the success of test-time search planners while requiring an order-of-magnitude fewer predictor invocations and $3$--$49\times$ lower wall-clock per decision. We report $98\%$ on PushT, $100\%$ on Reacher, $100\%$ on TwoRooms, and $92\%$ on the ogbench cube. The results show that, in a sufficiently regularized latent space, much of the structure that search discovers online can be amortized into a lightweight learned iterative policy, yielding fast, horizon-aware, nonlinear physical control without online optimization.
Chat is not available.
Successful Page Load