VidaForge: 비디오 사전학습 데이터 레시피를 위한 개방형 연구 인프라
VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes
September 6, 2026
저자: Yan Ma, Jiadi Su, Zhulin Hu, Ethan Chern, Linhao Zhang, Tiantian Mi, Pengfei Liu
cs.AI
초록
비디오 파운데이션 모델은 점점 더 대규모 사전학습 데이터에 의존하고 있지만, 그 이면의 종단 간 데이터 파이프라인은 대체로 폐쇄적이어서 검사하거나 재사용하기 어렵다. 비디오 데이터 레시피가 모델 사전학습에 미치는 영향을 이해하려는 연구자들은 초점을 맞춘 가설 하나를 검증하기 전에도 상당한 인프라를 구축해야 하는 경우가 많다. 우리는 VIDAFORGE를 제시한다. 이는 비디오 데이터 레시피를 원시 비디오부터 훈련 데이터셋까지 이어지는 실행 가능한 5단계 워크플로로 표현하는 개방형 연구 인프라이다. 이 워크플로에서 한 결정을 달리하면 모든 샘플이 어떻게 생성되었는지를 보존하면서 대체 데이터셋을 구성할 수 있다. 이 연구 워크플로를 시연하기 위해, 우리는 Wan 2.1과 V-JEPA 2.1의 초기 from-scratch 사전학습에서 서로 다른 커버리지와 품질의 데이터 레시피를 비교한다. 두 학습 목표 모두에서 더 넓은 커버리지를 가진 레시피가 가장 높은 다운스트림 벤치마크 점수를 달성한 반면, 손실 기반 평가는 서로 다른 레시피를 선호했다. 이 연구는 VidaForge가 데이터 레시피 선택을 다운스트림 모델 성능과 어떻게 연결하는지 보여준다. 우리는 또한 비디오 데이터 레시피 연구를 위한 세밀한 주석과 큐레이션 신호를 포함하며, 총 6,475시간에 달하는 314만 개의 장면 수준 클립을 담은 VIDAFORGE-3M을 공개한다.
English
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.