TKCAM: Text and Keyframe to Camera Trajectory Generation
Abstract
Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. To overcome the instability of direct 3D regression, we formulate camera dynamics as continuous 12-DoF kinematic sequences and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for robust multi-modal conditioning. A central feature of our framework is animator-style keyframe control: users can provide free-form text prompts alongside sparse key poses, and TKCAM seamlessly inpaints temporally coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a comprehensive cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM significantly surpasses recent state-of-the-art baselines on Fréchet Inception Distance (FID), text-motion matching scores, and retrieval metrics (R@K), producing cinematographically plausible motions while enabling precise spatial guidance. Our code, models, and dataset will be open-sourced upon acceptance.