PISCO: Precise Video Instance Insertion with Sparse Control
Abstract
AI video generation is moving beyond general generation, which relies on exhaustive prompt-engineering and "cherry-picking", towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, the core requirement is the ability to perform precise, targeted modifications. A key task is video instance insertion, which requires precise spatial-temporal placement, physically consistent scene interaction (e.g., shadows and reflections), and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To stabilize generation under sparse conditioning, we introduce Variable-Information Guidance and Distribution-Preserving Temporal Masking, complemented by geometry-aware conditioning. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms existing baselines and scales effectively as additional control signals are provided.