VidaForge:用於視訊預訓練資料配方之開放研究基礎設施
VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
September 6, 2026
作者: Yan Ma, Jiadi Su, Zhulin Hu, Ethan Chern, Linhao Zhang, Tiantian Mi, Pengfei Liu
cs.AI
摘要
影片基礎模型日益依賴大規模預訓練資料,然而其背後的端到端資料管線大多仍封閉,且難以檢視或重複使用。研究人員若想了解影片資料配方如何影響模型預訓練,往往需要先建立大量基礎設施,才能測試即使只是聚焦的假設。我們提出 VIDAFORGE,一套開放研究基礎設施,將影片資料配方表示為從原始影片到訓練資料集的可執行五階段工作流程。此工作流程中的一項決策可以改變,以建構替代資料集,同時保留每個樣本是如何產生的。為了展示此研究流程,我們在 Wan 2.1 與 V-JEPA 2.1 的早期從頭預訓練中,比較具有不同覆蓋範圍與品質的資料配方。在兩種學習目標上,覆蓋範圍較廣的配方取得最高的下游基準分數,而基於損失的評估則偏好不同的配方。本研究展示 VIDAFORGE 如何將資料配方選擇與下游模型效能連結起來。我們進一步釋出 VIDAFORGE-3M,包含 314 萬個場景級片段,總計 6,475 小時,並附有細粒度標註與策展訊號,供影片資料配方研究使用。
English
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.