ChatPaper.aiChatPaper

LAION-BVD:一個一千萬小時的開放影片資料集,用於多模態預訓練

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

August 25, 2026
作者: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
cs.AI

摘要

我們提出 LAION-BVD,一個用於多模態學習的大規模開放影片資料集,其中包含從 CommonCrawl 收集的 13 億個平台特定影片 URL。我們從中下載了 8,000 萬支影片,總時長達 1,000 萬小時。該資料集專為跨影片、音訊與影像模態的多模態預訓練而設計。利用內容感知場景偵測,我們提取片段,並針對這些片段合成生成影片與音訊的文字描述。在這些資料上訓練的模型於標準的影片-文字及音訊-文字基準測試中取得具有競爭力的表現,且隨著訓練規模或模型規模的增加呈現一致的改進。此外,我們透過提取場景變換幀,探索將影片幀作為影像-文字資料的替代來源。這些幀展現出與標準網路影像語料庫不同的視覺分佈,而在該資料集上訓練的模型亦達到優異的影像-文字檢索效能。我們將 LAION-BVD 釋出予研究社群,以前所未有的規模大幅擴展多模態影片的開放取用。
English
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.