LAION-BVD: 멀티모달 사전학습을 위한 1,000만 시간 규모의 공개 비디오 데이터셋
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
August 25, 2026
저자: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
cs.AI
초록
우리는 다중 모달 학습을 위한 대규모 개방형 비디오 데이터셋인 LAION-BVD를 제시한다. 이 데이터셋은 CommonCrawl에서 수집한 13억 개의 플랫폼별 비디오 URL을 포함하며, 이 중에서 총 1,000만 시간 분량의 8천만 개 비디오를 다운로드했다. 이 데이터셋은 비디오, 오디오, 이미지 모달리티에 걸친 다중 모달 사전 학습을 위해 설계되었다. 콘텐츠 인식 장면 검출을 사용하여 클립을 추출하고, 해당 클립에 대해 비디오 및 오디오 캡션을 합성 생성한다. 이 데이터로 학습된 모델은 표준 비디오-텍스트 및 오디오-텍스트 벤치마크에서 경쟁력 있는 성능을 달성하며, 학습 규모나 모델 규모가 증가함에 따라 일관된 성능 향상을 보인다. 또한, 장면 전환 프레임을 추출하여 이미지-텍스트 데이터의 대체 소스로서 비디오 프레임을 탐구한다. 이러한 프레임은 표준 웹 이미지 말뭉치와는 구별되는 시각적 분포를 나타내며, 이 데이터셋으로 학습된 모델은 강력한 이미지-텍스트 검색 성능을 달성한다. 우리는 LAION-BVD를 연구 커뮤니티에 공개한다. 이는 전례 없는 규모로 다중 모달 비디오에 대한 개방적 접근을 크게 확장한다.
English
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.