IDM-JEPA: Semi-Supervised Action Inference via Inverse Dynamics in Joint-Embedding Predictive Architectures
Abstract
Joint Embedding Predictive Architectures (JEPAs) learn world models directly in compact representation spaces, but existing methods depend on dense, frame-by- frame action supervision and must discard unlabeled transitions. We introduce IDM-JEPA, a semi-supervised framework that enables stable end-to-end JEPA training on partially labeled data by inferring missing controls with an Inverse Dynamics Model. Routing inferred latent action codes through the same encoder as ground-truth actions makes the two interchangeable during training and down- stream Model Predictive Control, at no cost to the planner, which searches over recorded actions throughout. On the continuous PushT benchmark, IDM-JEPA reaches 48.4% planning success from 5% action supervision, against 24.4% for an action-conditioned baseline trained on the labeled fraction alone under a matched budget of gradient steps. This correction is consequential: matching epochs gives the baseline only a fraction of updates and understates its success by as much as 36 points. After compute matching, the advantage is confined to extreme scarcity: at an equal step budget the baseline reaches 72.4% at 25% supervision and approaches full-data performance by 50%, leaving little headroom on a dataset of expert demonstrations whose unlabeled transitions are largely redundant with the labeled ones. Enforcing action-sensitivity through cycle-consistency on imagined transitions does not substitute for grounding codes in recorded ones. What deter- mines whether inferred actions substitute for labels is not the label fraction but whether the unlabeled pool adds state coverage; on expert demonstrations it largely does not.