SegmentWeave: Post-Processing framework for Enhanced Dense Video Captioning
Abstract
Dense video captioning (DVC) aims to localize events and generate a caption for each event in videos. For efficiency, many DVC pipelines rely on sparse fixed-interval frame sampling, which limits their ability to capture fine-grained semantic transitions and subtle visual details within each event. As a result, it is non-trivial for a trained model alone to incorporate visual information not present in the sampled frames, making further improvement difficult without retraining. We propose SegmentWeave, a training-free post-processing framework for event caption refinement. It refines captions independently of the training-time sampling rate and preserves the original caption as a semantic anchor. SegmentWeave revisits each predicted event through dense frame sampling, organizes it into semantically coherent micro-events, and augments the caption with visual details missed under sparse sampling. Experiments show that SegmentWeave consistently improves CIDEr, METEOR, and ROUGE-L across multiple pretrained DVC models as a drop-in module.