Think Before You Generate: Active Panoramic Exploration for Text-to-3D Scenes
Abstract
Creating explorable 3D environments from text is important for immersive content creation, simulation, and interactive design. Recent methods often rely on image or video generative priors by synthesizing views along fixed or manually specified trajectories and lifting them into 3D. However, such passive pipelines cannot adapt view generation to the evolving scene memory, often leaving disoccluded regions incomplete and producing holes, floaters, and inconsistent geometry under free-viewpoint exploration. We argue that the key challenge is not to generate more views, but to decide where to generate useful views according to the current scene state. To this end, we propose ActPano3D, an active text-to-3D scene generation framework based on memory-guided panoramic exploration. ActPano3D formulates trajectory selection as spatial deliberation over the evolving scene memory, where candidate actions are evaluated by their expected utility for exploring unknown regions, repairing uncertain geometry, and improving the future 3D world model. The framework integrates expansion-refinement active planning, memory-routed panoramic generation, and reliability-aware scene memory fusion into a closed-loop generation-and-reconstruction process. The accumulated memory is further optimized into a renderable 3D Gaussian Splatting scene for free-viewpoint exploration. Experiments show that ActPano3D improves scene coverage, reduces rendering holes, and achieves better rendering quality and text alignment than existing text-to-3D scene generation baselines.