ChatPaper.aiChatPaper

VidaForge:動画事前学習データレシピのためのオープン研究基盤

VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

September 6, 2026
著者: Yan Ma, Jiadi Su, Zhulin Hu, Ethan Chern, Linhao Zhang, Tiantian Mi, Pengfei Liu
cs.AI

要旨

動画基盤モデルは大規模な事前学習データにますます依存するようになっているが、その背後にあるエンドツーエンドのデータパイプラインは依然として大部分が非公開であり、検査や再利用が難しい。動画データレシピがモデル事前学習に与える影響を理解しようとする研究者は、焦点を絞った仮説を検証する前でさえ、しばしば相当なインフラストラクチャを構築する必要がある。我々は、動画データレシピを、生動画から学習データセットに至る実行可能な5段階ワークフローとして表現するオープン研究基盤 VIDAFORGE を提示する。このワークフローにおける決定を変化させることで、すべてのサンプルがどのように生成されたかを保持したまま、代替データセットを構築できる。この研究ワークフローを実証するため、Wan 2.1 と V-JEPA 2.1 のスクラッチからの初期事前学習において、カバレッジと品質の異なるデータレシピを比較する。両方の学習目的において、より広いカバレッジのレシピが下流ベンチマークで最高スコアを達成する一方、損失ベースの評価では異なるレシピが有利となる。本研究は、VidaForge がデータレシピの選択を下流のモデル性能にどのように結び付けるかを示す。さらに我々は、動画データレシピ研究のための細粒度アノテーションとキュレーション信号を備え、合計 6,475 時間に及ぶ 314 万件のシーンレベルのクリップを含む VIDAFORGE-3M を公開する。
English
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.