可解释的感知言语脑磁图解码:皮层源及驱动检索的刺激特征
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
August 2, 2026
作者: Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
cs.AI
摘要
通过以CLIP风格目标函数针对wav2vec 2.0音频嵌入训练的深度网络,可以从非侵入性脑磁图(MEG)记录中检索到感知语音的短片段。然而,其权重并不映射到电生理量上,目前仍不清楚哪些语音属性驱动了检索。
我们基于一个高性能的MEG到音频检索架构,但重新设计了其前端和解码器。其空间注意力作用于扁平传感器布局;我们将其替换为定义在三维MEG头盔几何结构上的球谐函数。我们将受试者特定表示从270个分支减少到25个,为每个分支添加时间滤波器以在空间和时间上匹配神经元源,并使卷积解码器变得更浅。在训练前移除眼动和心电成分,以降低刺激锁定捷径的风险。
在MEG-MASC上,该模型在六个训练模型的1005个候选中达到39.75 ± 0.34%的前1准确率,而解码器参数减少了约20倍。其权重映射到源空间,恢复了与语音感知网络一致的生成源,而左侧化分支携带了右侧不明显的高频节律成分。配对MEG遮挡显示,19个刺激特征中有15个有贡献,其中静音、声音强度、元音和声学起始的影响最大。随机词表则表现相反:将叙事性MEG替换到其中可改善检索,这表明没有叙事结构的活动所携带的可恢复信息少于连贯语音期间的活动。wav2vec目标可以减少到约十二个学习特征维度而不损失准确率,而强烈的时间压缩则导致明显损失。
总之,源映射和输入干预揭示了驱动检索的因素。
English
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval.
We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts.
On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss.
Together, source mapping and input interventions reveal what drives retrieval.