Start-Controlled Evaluation of JEPA Controllers: Test-Time Search Collapses on Manipulation, and Amortization Needs a Horizon It Should Infer
Abstract
Forming behavioral policies from pretrained JEPA representations requires a test-time procedure for sequential decision making, typically sampling-based search or model predictive control (MPC). Recent work claims that amortized goal-conditioned inverse dynamics (GC-IDM) offers a cheaper alternative that matches or beats this search. We propose a start-controlled evaluation protocol for JEPA-based controllers on the existing benchmark suite, and we identify and remove a deployment-critical assumption the amortized policy relies on. Start-controlled evaluation measures from the true episode start rather than a random mid-episode state. Our key finding is that test-time search collapses on the manipulation task under the standard latent-distance planning cost. Random-start evaluation instead inflates success rates unevenly across tasks and can even reverse in sign. We then show that the amortized policy's performance depends on privileged information. GC-IDM is handed the ground-truth temporal distance to the target to guide action generation, information a deployed system rarely has (e.g.\ when goals come from a hierarchical sub-goal generator). We propose VC-IDM, which removes this dependence with a learned temporal-distance value, where a deployed system operates. Across four environments and two representations, VC-IDM wins when the target distance is unknown, variable, or beyond the training range. Unlike GC-IDM, it improves with additional budget rather than degrading. Evaluation methodology and conditioning mechanisms on pretrained representations change the conclusions drawn.