Seeing the World through Any Eyes
Abstract
Egocentric video generation aims to synthesize first-person visual experiences, enabling applications in filmmaking, virtual reality, and embodied AI. Generating egocentric videos from exocentric observations is particularly challenging, as it requires reasoning across large viewpoint changes, limited visual overlap, and substantially different camera motions. To address these modeling challenges, we propose EgoEye, a framework that generates egocentric videos from a single exocentric input video. EgoEye integrates reward-guided egocentric context reasoning for inferring unseen first-person content, multi-stage motion alignment for enforcing cross-view temporal consistency, and egocentric pretraining for improving first-person realism. To support training and evaluation under diverse and distortion-free exo-ego settings, we further introduce EgoScape, a large-scale dataset comprising 19.5K time-aligned exo-ego video pairs and 24.2K in-the-wild egocentric videos. EgoScape provides paired cross-view supervision while enriching the first-person visual priors needed for realistic egocentric synthesis. Extensive experiments demonstrate that the proposed EgoEye generates realistic and temporally consistent egocentric videos across diverse scenarios.