Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Abstract
Grounded long-video question answering, or Grounded LVQA, is the task of answering a question about a long video while also locating the short time interval in the video that supports the answer. Recent agentic methods approach this problem as a multi-turn exploration process with a single cropvideo(start, end) action. This action allows the model to progressively narrow its search from coarse to fine regions, but it does not provide a direct way to backtrack from a fine-grained mistake to a broader context. As a result, these agents often stop after only two turns and are unable to recover once they descend into the wrong part of the video. We propose VideoTreeSearch (VTS), a framework that formulates grounded LVQA as an iterative, self-correcting search over an adaptive temporal tree. VTS builds a non-uniform tree from scene boundaries so that each node corresponds to a semantically coherent video segment. It then trains an agent to navigate this tree using four discrete actions: zoomin, zoom_out, shift, and answer. These actions make backtracking and recovery explicit and learnable, rather than leaving them as implicit behaviors. To train the agent, we introduce a trajectory synthesis pipeline that generates multi-step navigation paths through the tree, including intentional detours into incorrect branches followed by recovery. These trajectories are first used for supervised fine-tuning and then for reinforcement learning with rewards based on grounding quality and answer accuracy. On three Grounded LVQA benchmarks—CG-Bench, Haystack-LVBench, and Haystack-Ego4D—VTS outperforms the strongest previous agentic methods by 12.5 mIoU on CG-Bench and 7.4 T-F1 on Haystack-Ego4D. The learned policy also transfers to general long-video question answering, surpassing all prior agentic baselines on Video-MME, MLVU, and LVBench by up to 7.1 accuracy points. Ablation studies show that self-correcting hierarchical search is the key factor behind these improvements: removing either adaptive descent or explicit backtracking leads to substantial performance drops.