LibriBrain100:一百小时广泛而深入的MEG数据,用于大规模神经语音解码
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
August 25, 2026
作者: Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
cs.AI
摘要
我们介绍了LibriBrain100,这是一个大规模用于语音解码的MEG数据集,从设计之初就致力于可复现、标准化的评估。LibriBrain100的规模是原始LibriBrain发布版本的两倍以上,包含超过100小时的高质量MEG数据,这些数据是在受试者聆听自然连续语音时采集的。其中约80小时来自单个受试者,LibriBrain100在个体内深度神经数据方面创下了新纪录(是下一个可比数据集的8倍,约为其他数据集的80倍)。为了展示这种深度优先设计的价值,我们在一个词分类基准上进行了评估——这是通向无创脑到文本解码这一公开挑战的日益成熟的基石。使用现有解码模型,我们达到了最先进的性能——这既验证了录音质量,也验证了大规模个体内数据的价值。由于在现实应用中为每位用户收集80小时的数据并不实际,我们还从32名受试者中各收集了约40分钟的额外数据。使用相同的词分类基准,我们展示了广泛多受试者数据的价值:对预训练模型进行有监督微调可以大幅弥补单个受试者数据有限的问题。我们提供标准的训练集、验证集和测试集划分,所有这些都可以通过一个开源Python库复现,该库支持便捷下载、可选预处理以及为常见深度学习框架加载数据。此外,该数据集和评估基础设施将与一项开放机器学习竞赛一同发布,并配有公开排行榜,用于标准化基准测试。最终,我们希望LibriBrain100能够加速推进实用型无创脑机接口的发展,为严重瘫痪患者恢复沟通能力。
English
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With sim80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8times more than the next comparable dataset and roughly 80times more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected sim40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.