PRISM: Programming Interactive Scenes from Monocular Images for Embodied Simulation
Abstract
The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the physical world. However, transforming raw visual observations into simulation-ready scenes remains challenging due to the lack of physical grounding and scene-level interactivity in current image-to-URDF methods. We propose PRISM, a framework that reformulates monocular scene reconstruction as a procedural programming task for interactive 3D environments. By leveraging the zero-shot reasoning and code synthesis of MLLMs, PRISM translates a single RGB image into executable programs defining object geometry, articulation, and physical properties. To ensure simulation readiness, it incorporates a physics-in-the-loop mechanism that iteratively refines the generated programs by validating their execution in a physics engine. This feedback loop enforces physically plausible object articulations, valid object compositions and interactions, and accurate spatial relationships. Experiments demonstrate that PRISM significantly outperforms open-loop approaches and prior monocular reconstruction models. Notably, PRISM-generated scenes support complex downstream tasks such as stable stacking and fine-grained manipulation, which are difficult to achieve with existing methods.