Video Generation Research Needs a Legally Sustainable Data Infrastructure
Abstract
What happens if foundational training data in a research field is withdrawn? The video generation community has already faced this issue. That is, WebVid-10M, which supported over 30 video generation models, was taken down after Shutterstock issued a cease-and-desist order. A comparable situation may emerge soon, as Panda-70M is now involved in a class-action lawsuit filed in April 2026 against Apple, Amazon, and OpenAI. In addition, YouTube has made AI training opt-in disabled by default. Against this backdrop, we show why video generation research needs a legally sustainable data infrastructure urgently with three findings: (1) The risk. About 73.3% of major video datasets come from YouTube. Over 770 paper–dataset dependency instances exist. If just one dataset, like HD-VILA-100M, is withdrawn, it could affect five other related datasets. (2) The legal landscape is fragmenting. Recent legal and regulatory developments do not provide a stable permission boundary for AI video training: Thomson Reuters v. Ross rejected fair use in one non-generative AI setting, and other AI copyright rulings remain fact-specific. (3) The frontier remains unchanged. On June 19, 2025, Google told CNBC that its main video model, Veo 3, is trained on a subset of YouTube videos, and creators cannot opt out. Using just one percent of this library gives about forty times more training data than other models. Several creators said they were not told or asked about this use. In short, the video generation field, from older academic datasets to new developments, still depends on content whose licensing, creator-consent, and platform-access status remain unresolved. To solve this, we suggest a five-part research plan: focus on Creative Commons collections, gather public-domain content, build systems for creators to opt in, set standards for tracking content origins, and check data compliance at the publication level.