LibriBrain100: 大規模な神経音声デコーディングのための広範かつ詳細な100時間MEGデータ
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
August 25, 2026
著者: Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
cs.AI
要旨
LibriBrain100は、音声デコーディングのための大規模MEGデータセットであり、再現可能で標準化された評価のためにゼロから設計されている。LibriBrain100は元のLibriBrainリリースの規模を2倍以上に拡大し、被験者が自然な連続音声を聴取している間に収集された100時間以上の高品質MEGを提供する。単一被験者からの約80時間のデータにより、LibriBrain100は被験者内の深い神経データの新記録を樹立する(次に匹敵するデータセットの8倍、他のデータセットのおよそ80倍)。この深さ優先設計の効果を示すために、我々は単語分類ベンチマークで評価を行う。これは、非侵襲的な脳からテキストへのデコーディングという未解決の課題に向けた、ますます確立されつつある足掛かりである。既存のデコーディングモデルを用いて最先端の性能を達成し、記録の品質と大規模な被験者内データの価値の両方を検証する。実世界の応用ではユーザーごとに80時間のデータを収集することは現実的ではないため、32名の被験者それぞれから約40分の追加データも収集した。同じ単語分類ベンチマークを用いて、広範な多被験者データの価値を示す。事前学習モデルの教師ありファインチューニングは、被験者ごとの限られたデータを大幅に補うことができる。我々は標準的な訓練、検証、テスト分割を提供し、これらはすべてオープンソースのPythonライブラリを通じて再現可能である。このライブラリは、簡単なダウンロード、任意の前処理、一般的な深層学習フレームワーク向けのデータ読み込みをサポートする。さらに、データセットと評価インフラストラクチャは、標準化されたベンチマーキングのための公開リーダーボードを備えたオープンな機械学習コンペティションとともに公開されている。最終的に、LibriBrain100が、重度の麻痺を抱える人々にコミュニケーションを回復させることができる実用的な非侵襲的ブレイン・コンピュータ・インターフェースへの進歩を加速することを期待している。
English
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With sim80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8times more than the next comparable dataset and roughly 80times more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected sim40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.