ChatPaper.aiChatPaper

LAION-BVD:マルチモーダル事前学習のための1,000万時間のオープンビデオデータセット

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

August 25, 2026
著者: Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
cs.AI

要旨

本稿では、マルチモーダル学習のための大規模オープンビデオデータセットであるLAION-BVDを提案する。このデータセットは、CommonCrawlから収集された13億個のプラットフォーム固有のビデオURLを含み、そのうち8000万本のビデオ(総再生時間1000万時間)をダウンロードした。本データセットは、ビデオ、オーディオ、画像の各モダリティにわたるマルチモーダル事前学習用に設計されている。内容認識型シーン検出を用いてクリップを抽出し、そのクリップに対してビデオおよびオーディオのキャプションを合成的に生成する。これらのデータで学習されたモデルは、標準的なビデオ-テキストおよびオーディオ-テキストのベンチマークにおいて競争力のある性能を達成し、学習規模またはモデル規模の増加に伴って一貫した改善が見られる。さらに、シーン変化フレームを抽出することで、ビデオフレームを画像-テキストデータの代替ソースとして探求する。これらのフレームは、標準的なウェブ画像コーパスとは異なる視覚的分布を示し、本データセットで学習されたモデルは強力な画像-テキスト検索性能を達成する。我々はLAION-BVDを研究コミュニティに公開する。これにより、前例のない規模でマルチモーダルビデオへのオープンアクセスが大幅に拡大される。
English
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.