Reconstruction vs. Prediction Pre-Training of Discrete Video Representations for Continuous Control
Abstract
Self-supervised video encoders are commonly trained to either reconstruct pixels or predict latent features, yet it remains unclear how these objectives shape representation for downstream dynamics. We present a capacity-matched study isolating this choice: two identical video encoders with a shared FSQ bottleneck are trained under autoencoding and V-JEPA objectives on a continuous control environment. We introduce a simple asymmetric discretization scheme to apply the same bottleneck to the predictive model. The resulting representations exhibit a clear trade-off: the autoencoder preserves substantially more physical state information and achieves slightly better one-step latent prediction, while the predictive encoder accumulates less error under autoregressive rollout. Structural probes suggest that this advantage is associated with greater robustness to low-level visual perturbations and stronger dependence on temporal context. These results indicate that predictive and reconstructive objectives induce distinct representations of dynamics, trading state fidelity for properties that may benefit long-horizon prediction.