LibriBrain100: 대규모 신경 음성 디코딩을 위한 100시간의 광범위하고 심층적인 MEG 데이터
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
August 25, 2026
저자: Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones
cs.AI
초록
우리는 LibriBrain100을 소개한다. 이는 처음부터 재현 가능하고 표준화된 평가를 위해 설계된 음성 디코딩용 대규모 MEG 데이터셋이다. LibriBrain100은 기존 LibriBrain 릴리스보다 두 배 이상 큰 규모로, 피험자들이 자연스러운 연속 음성을 듣는 동안 수집된 100시간 이상의 고품질 MEG 데이터를 포함한다. 단일 피험자로부터 약 80시간을 확보함으로써, LibriBrain100은 심층적인 피험자 내 신경 데이터의 새로운 기록을 세웠다(다음으로 유사한 데이터셋보다 8배, 다른 데이터셋보다 약 80배 더 많음). 이러한 심층 우선 설계의 효과를 입증하기 위해, 우리는 단어 분류 벤치마크에서 평가를 수행한다. 이는 비침습적 뇌-텍스트 디코딩이라는 공개 과제로 가는 점점 더 확립된 디딤돌이다. 기존 디코딩 모델을 사용하여 우리는 최첨단 성능을 달성했으며, 이는 녹음 품질과 대규모 피험자 내 데이터의 가치를 모두 검증한다. 사용자당 80시간의 데이터 수집이 실제 응용에서 비현실적이기 때문에, 우리는 또한 32명의 각 피험자로부터 약 40분의 추가 데이터를 수집했다. 동일한 단어 분류 벤치마크를 사용하여, 우리는 광범위한 다중 피험자 데이터의 가치를 입증한다: 사전 훈련된 모델의 지도 파인튜닝은 제한된 피험자별 데이터를 상당 부분 보완할 수 있다. 우리는 표준 훈련, 검증, 테스트 분할을 제공하며, 모든 것은 쉬운 다운로드, 선택적 전처리, 일반적인 딥러닝 프레임워크용 데이터 로딩을 지원하는 오픈소스 파이썬 라이브러리를 통해 재현 가능하다. 또한 데이터셋과 평가 인프라는 공개 머신러닝 대회와 함께 공개되어 표준화된 벤치마킹을 위한 공개 리더보드를 제공한다. 궁극적으로, 우리의 바람은 LibriBrain100이 심각한 마비로 고통받는 사람들의 의사소통을 회복할 수 있는 실용적인 비침습적 뇌-컴퓨터 인터페이스 쪽으로의 진전을 가속화하는 것이다.
English
We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired while subjects listened to naturalistic continuous speech. With sim80 hours from a single subject, LibriBrain100 sets a new record for deep, within-subject neural data (8times more than the next comparable dataset and roughly 80times more than other datasets). To demonstrate the payoff of this depth-first design, we evaluate on a word-classification benchmark---an increasingly well-established stepping stone towards the open challenge of noninvasive brain-to-text decoding. Using an existing decoding model, we achieve state-of-the-art performance---validating both the quality of the recordings and the value of within-subject data at scale. Because collecting 80 hours of data per user is impractical for real-world applications, we also collected sim40 minutes of additional data from each of 32 subjects. Using the same word-classification benchmark, we demonstrate the value of broad multi-subject data: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data. We provide standard train, validation, and test splits, all reproducible through an open-sourced Python library that supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks. In addition, the dataset and evaluation infrastructure are being released alongside an open machine-learning competition with a public leaderboard for standardised benchmarking. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non-invasive brain-computer interfaces, capable of restoring communication to people living with severe paralysis.