LibriBrain100:一百小時廣泛且深入的腦磁圖數據,用於大規模神經語音解碼
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
August 25, 2026
作者: Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
cs.AI
摘要
我們介紹 LibriBrain100,這是一個大規模的腦磁圖(MEG)語音解碼資料集,從根本上是為了可重現、標準化的評估而設計。LibriBrain100 將原始 LibriBrain 版本的規模擴大一倍以上,提供超過 100 小時的高品質腦磁圖資料,內容為受試者在聆聽自然連續語音時所記錄的訊號。憑藉來自單一受試者的約 80 小時資料,LibriBrain100 創下了受試者內神經資料深度的新紀錄(比下一個可比資料集多 8 倍,比其他資料集大約多 80 倍)。為了展示這種深度優先設計的價值,我們在詞彙分類基準上進行評估——這是朝向非侵入式腦到文本解碼這項開放挑戰日益成熟的踏腳石。我們使用現有的解碼模型,取得了最先進的效能——同時驗證了錄音品質以及大規模受試者內資料的價值。由於在實際應用中,每位使用者收集 80 小時的資料並不切實際,我們也從 32 位受試者每人收集了約 40 分鐘的額外資料。我們使用相同的詞彙分類基準,展示了廣泛的多受試者資料的價值:對預訓練模型進行監督式微調,可以大幅彌補每位受試者資料有限的問題。我們提供標準的訓練、驗證和測試分割,並可透過一個開源的 Python 函式庫完整重現,該函式庫支援輕鬆下載、可選的預處理,以及供常見深度學習框架使用的資料載入。此外,該資料集和評估基礎設施將與一場開放的機器學習競賽一同發布,並附有公開排行榜以供標準化基準比較。最終,我們希望 LibriBrain100 能加速實用非侵入式腦機介面的發展,使其有能力為嚴重癱瘓患者恢復溝通能力。
English
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With sim80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8times more than the next comparable dataset and roughly 80times more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected sim40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.