知覚された音声の解釈可能なMEGデコーディング:検索を駆動する皮質源と刺激特徴
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
August 2, 2026
著者: Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
cs.AI
要旨
短い音声知覚区間は、wav2vec 2.0オーディオ埋め込みに対するCLIPスタイルの目的関数を用いて訓練された深層ネットワークにより、非侵襲的な脳磁図(MEG)記録から検索できる。しかし、その重みは電気生理学的量に直接対応しておらず、どの音声特性が検索を駆動しているかは依然として不明である。
我々は、高性能なMEG-to-audio検索アーキテクチャを基盤としつつ、そのフロントエンドとデコーダの両方を再設計する。その空間的注意は平坦化されたセンサー配置に作用するが、我々はこれを3次元MEGヘルメット形状上で定義される球面調和関数に置き換える。被験者固有の表現を270ブランチから25ブランチに削減し、各ブランチに時間フィルタを追加して空間的・時間的に神経源と対応させ、畳み込みデコーダをより浅い構造にする。訓練前に眼球性および心臓性の成分を除去し、刺激同期型の近道学習のリスクを低減する。
MEG-MASCにおいて、本モデルは6つの訓練済み解にわたり、1005候補中のTop-1精度39.75±0.34%を達成し、デコーダのパラメータ数は約20分の1である。その重みはソース空間に写像され、音声知覚ネットワークと整合する生成源を回復する一方、左半球優位のブランチは右側では明瞭でない高周波数の律動的成分を担う。ペアMEGオクルージョンにより、19の刺激特徴のうち15が寄与し、無音、音強度、母音、および音響開始が最も大きな効果を示す。ランダムな単語リストでは逆の挙動が見られる:ナラティブMEGをそれらに置換すると検索が改善され、ナラティブ構造を伴わない活動は、一貫した音声中の活動よりも回復可能な情報が少ないことを示す。wav2vecターゲットは、精度を損なうことなく約12の学習特徴次元に削減できるが、強い時間的圧縮は明確な損失を引き起こす。
総合すると、ソースマッピングと入力介入は、何が検索を駆動するかを明らかにする。
English
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval.
We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts.
On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss.
Together, source mapping and input interventions reveal what drives retrieval.