可解釋的感知語音腦磁圖解碼:驅動檢索的皮質來源與刺激特徵
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
August 2, 2026
作者: Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
cs.AI
摘要
感知語音的短片段可透過深度網路從非侵入性腦磁圖(MEG)記錄中檢索出來;這些深度網路以 CLIP 風格目標函數針對 wav2vec 2.0 音訊嵌入進行訓練。然而,其權重無法對應至電生理量,且目前仍不清楚哪些語音特性驅動了檢索。
我們基於一個高性能的 MEG 到音訊檢索架構,但重新設計了其前端與解碼器。原有的空間注意力作用於展平後的感測器佈局;我們以定義在三維 MEG 頭盔幾何結構上的球諧函數取代之。我們將受試者特定表徵從 270 個分支減少至 25 個分支,並為每個分支加入時間濾波器,使其在空間與時間上對應到神經元來源,同時將卷積解碼器改為更淺的結構。訓練前移除眼動與心臟成分,以降低依賴刺激鎖定捷徑的風險。
在 MEG-MASC 上,六個訓練模型中,該模型在 1005 個候選項目中達到 39.75 ± 0.34% 的 Top-1 準確率,且解碼器參數約為原本的二十分之一。其權重可映射至源空間,恢復出與語音感知網路一致的產生源;左側化的分支攜帶較高頻率的節律成分,而右側並不明顯。成對 MEG 遮蔽實驗顯示,19 個刺激特徵中有 15 個具貢獻,其中靜音、聲音強度、母音與聲學起始的效應最大。隨機詞表則表現相反:將敘述性 MEG 替換進去可改善檢索,這表示缺乏敘事結構的活動所攜帶的可恢復資訊,少於連貫語音期間的活動。wav2vec 目標可縮減至約十二個學習特徵維度而不損失準確率,但強烈時間壓縮則會造成明顯損失。
綜合而言,源映射與輸入干預揭示了檢索的驅動因素。
English
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval.
We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts.
On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss.
Together, source mapping and input interventions reveal what drives retrieval.