Slide P2V-Bench: A Cross-Domain Benchmark for Slide-Centric Scientific Paper-to-Presentation Video Generation
Abstract
Scientific communication increasingly relies on slide-based narrated videos rather than static PDFs. However, transforming a paper into a presentation video is not a conventional text-to-video task. It requires preserving technical rigor, rendering dense multimodal assets faithfully, and synchronizing narration, subtitles, and cursor motion over long horizons. Current research faces two critical bottlenecks: (1) Benchmark Gap: existing evaluations are restricted to narrow domains or structured LaTex inputs with limited long-range supervision; (2) Evaluation Gap: current protocols often overlook the joint measurement of content fidelity, channel alignment, and actual knowledge transfer. To bridge these gaps, we introduce SLIDE P2V-BENCH, a cross-domain bench mark comprising 108 high-quality paper-video pairs spanning four domains and various durations, coupled with an auditable multi-dimensional evaluation suite. Furthermore, we propose PAPER2MEDIA, a slide-grounded agentic pipeline that produces editable presentations by synchronizing decoupled media channels via a shared semantic clock. Extensive experiments demonstrate that PAPER2MEDIA significantly outperforms existing automated baselines and, notably, achieves performance comparable to human-authored presentations in terms of narrative coherence, visual quality, and knowledge conveyance efficiency.