When Does Hierarchical Abstraction in World Models Actually Learn Anything? A Controlled Diagnostic Study
Regina Nam
Abstract
Hierarchical predictive architectures such as H-JEPA (LeCun, 2022) propose learning abstractions at multiple temporal and spatial scales, but the proposal is architectural rather than algorithmic: it specifies neither a training objective for the abstract level nor a way to verify that a trained abstraction encodes anything a random projection of the same shape would not. We introduce a diagnostic protocol built around a random-abstractor control — comparing a trained top level against an identically-parameterized but untrained one — and apply it across a $2 \times 2$ design crossing observability (full vs. egocentric) with abstractor type (instantaneous vs. recurrent) in controlled gridworld environments with exact ground truth at both scales. Three of four conditions yield abstractions that are statistically indistinguishable from random projections (or slightly worse), for two distinct reasons: under full observability the abstraction has nothing to add, since the slow variable is already linearly decodable from the fast level; under partial observability an instantaneous abstractor lacks any mechanism to integrate the history required to recover it. Only the combination of partial observability and a recurrent abstractor produces a measurably learned abstraction, replicated across 3 seeds (top-level probe accuracy gap over random: $19.91 \pm 3.36$pp, vs. $0.83 \pm 0.17$pp under full observability). We further report a negative result: attempts to establish a quantitative relationship between the density of persistent visual landmarks and abstraction quality in an open-world setting did not survive multi-seed replication, despite a promising single-run trend. Our results give an empirical account of why architectures such as DreamerV3's RSSM (Hafner et al., 2023) combine a recurrent deterministic state with an instantaneous latent, and provide a reusable diagnostic that we argue should accompany any claim that a hierarchical model has learned a meaningful abstraction.
Chat is not available.
Successful Page Load