GLIMPSE: Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
Abstract
Text carries much of the actionable information in everyday scenes, but recognizing it from a wearable is constrained by power: OCR needs high-resolution input, and streaming high-resolution video drains the battery. We observe that this constraint is asymmetric. Comparing our system configurations at a fixed capture rate, coarsening the video a reasoning model sees costs one accuracy point, while coarsening the input OCR sees costs 32. GLIMPSE exploits that asymmetry in a continuous wearable stream: selective full-resolution OCR runs on-device and only transcribed text is transmitted, while low-resolution video is streamed for visual context. A three-stage frame selector cuts OCR invocations by 67.7\%, and an OCR Session Manager reassembles the resulting sparse, irregularly-timed payloads into a temporally aligned context state. On 208 text-VQA questions over 5.6 hours of egocentric video, the partition recovers 31 accuracy points over an equal-bandwidth baseline (41\% to 72\%) for 0.02x additional device power.