Placing the Bottleneck: Spatial Latent Scene Representations for World Models
Maciej Lewandowski ⋅ Michael Pearce ⋅ Tsun-Yi Yang ⋅ Thomas J Cashman
Abstract
Reliable and useful models of the physical world must preserve geometric structure as they generate observations across viewpoints. We explore scene representations that can integrate observations and be queried repeatedly under new viewpoints to give geometrically consistent predictions. Highway Novel View Synthesis architectures preserve observation-derived features and their camera provenance, achieving strong reconstruction quality, but their representation and decoder-visible context grow with the number of observations. Fixed bottlenecks instead compress observations into a bounded, target-independent scene memory, making them attractive as compact and reusable world representations, but often at a cost in reconstruction quality. We investigate whether this performance gap between highway and bottleneck approaches is caused only by information compression, or also by the loss of an explicit geometric relation between the scene representation and the renderer. In a scaled-down bottleneck SVSM experiment, increasing the number of unstructured scene tokens produces only modest gains and does not close the gap to a separately trained highway reference. We then propose a spatially placed variant of a SceneTok-style bottleneck architecture in which each latent slot is assigned a 3D region. A shared, parameter-free relation between patch-centre rays and latent cells guides both context-to-memory writes and memory-to-target reads. On the CLEVR-TR dataset, spatial placement makes additional memory capacity substantially more useful; in the larger model with $K=500$ scene tokens, spatial placement improves mean-scene PSNR by 3.65 dB, averaged over three matched training seeds. The resulting memory supports localized scene edits: masking geometry-selected groups of cells produces spatially concentrated changes in the rendered scene, giving the latent memory a spatially interpretable structure. We validate this approach on MultiShapeNet, and find that the spatial variant improves all six matched model-scale and memory-capacity configurations by $0.70$--$1.08$ dB. Together, these results suggest that the effectiveness of compact scene memory depends not only on capacity (i.e.~how much information it can store), but also on whether the scene representation and renderer share a stable geometric relation.
Chat is not available.
Successful Page Load