Infinite Token Streaming via Time-based Voxel Grid for Visual Geometry Grounded Transformers
Abstract
Visual Geometric Grounded Transformers have emerged as a powerful paradigm for 3D scene construction, replacing classical approaches in the structure from motion problem. However, deploying such systems in continuous and long-horizon streaming scenarios introduces severe computational and memory bottlenecks, as the token cache scales linearly with the stream duration. Meanwhile, existing token pruning and eviction policies mitigate memory footprint at the expense of forgetting, discarding critical geometric context from previously observed viewpoints. In this work, we propose a framework for Infinite Token Streaming via Time-based Voxel Grid, where we introduce a continuous spatial hashmap managed by a time-based cache that dynamically aggregates discarded tokens, which share the same voxel (or world) coordinates, into prototypes via online density-weighted updates while evicting the least recently observed prototypes to enforce a bounded memory footprint, demonstrating performance improvements over our baselines.