PercepCap: Video Captioner with Structured Spatio-Temporal Perception
Abstract
Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing Multi-modal Large Language Models (MLLMs) usually generate captions directly from video inputs without exposing the perceptual evidence behind their descriptions. As a result, object, event, or temporal mistakes are only observed in the final text, making it difficult to identify the underlying perceptual errors or optimize them directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evidence explicit before producing the final caption. Specifically, PercepCap follows a perceive-describe generation chain, where the model first produces a spatio-temporal perception trace comprising object trajectories and temporal events according to the video, and then generates the final caption conditioned on the perceived evidence. To support this new generation paradigm, we design a two-stage training strategy for PercepCap. Perceive-then-Describe Supervised Fine-tuning (PD-SFT) adapts the model from caption-only generation to the proposed perceive-describe chain, while Perception-Grounded Reinforcement Learning (PG-RL) further optimizes perception trace and caption quality with joint rewards over object tracking, temporal event, and object/action description coverage. Caption-Anchored Perception Data Construction builds this supervision by first generating a caption-only description, extracting the objects and events it mentions, and grounding them back in the video with boxes and timestamps. This yields caption-aligned perception data that provides both SFT supervision and RL references, ensuring that the explicit perception trace and final caption describe the same objects and events. Across direct caption evaluation such as DREAM-1K, CaReBench and VidCapBench and caption-to-QA evaluation like ShortVidBench and MotionBench. PercepCap consistently improves the Qwen3-VL baseline and achieves leading open-source caption quality.