UTOPI: Efficient Egocentric Long-Video Understanding in AR via User-Guided Token Pre-Compression
ziqi wang ⋅ Su Chen ⋅ Qiance Tang ⋅ Jieyu Lin ⋅ Ziyun Li ⋅ Barbara De Salvo ⋅ Sai Qian Zhang
Abstract
As long-video understanding becomes increasingly important, augmented reality (AR) devices offer a natural platform for deploying egocentric video intelligence. Yet first-person videos often contain substantial temporal redundancy, creating heavy memory and compute demands for resource-limited AR hardware. We present~\textit{UTOPI}, a plug-and-play pre-compression module tailored to egocentric long-video processing on AR devices. UTOPI first uses user motion cues to estimate cross-frame overlap and remove redundant content through a viewpoint-shift-robust pruning strategy. It then leverages eye-tracking signals to detect user fixation regions, separating attention-relevant foreground from background and preserving important content through a fine-tuned neural adjustment module. Evaluated across multiple benchmarks and model backbones, UTOPI reduces video tokens by up to $95%$ while maintaining task performance.
Chat is not available.
Successful Page Load