ChatPaper.aiChatPaper

LAION-BVD:一个面向多模态预训练的一千万小时开放视频数据集

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

August 25, 2026
作者: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
cs.AI

摘要

我们提出了 LAION-BVD,一个用于多模态学习的大规模开放视频数据集,其中包含从 CommonCrawl 收集的 13 亿个特定平台的视频 URL。从中,我们下载了 8000 万个视频,总时长达 1000 万小时。该数据集旨在跨视频、音频和图像模态进行多模态预训练。利用内容感知的场景检测,我们提取片段,并为其合成生成视频和音频描述。在这些数据上训练的模型在标准视频-文本和音频-文本基准上取得了具有竞争力的性能,并且随着训练规模或模型规模的增加而持续提升。此外,我们通过提取场景变化帧,探索将视频帧作为图像-文本数据的替代来源。这些帧展现出与标准网络图像语料库不同的视觉分布,基于该数据集训练的模型取得了强大的图像-文本检索性能。我们将 LAION-BVD 发布给研究社区。它以前所未有的规模显著扩大了对多模态视频的开放访问。
English
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.