ChatPaper.aiChatPaper

지각된 말소리의 해석 가능한 MEG 디코딩: 피질 소스와 복원을 주도하는 자극 특성

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

August 2, 2026
저자: Ilia Semenkov, Daria Kleeva, Ivan Dakhtin, Zarina Maksudova, Alex Ossadtchi
cs.AI

초록

인지된 음성의 짧은 구간은 CLIP 방식의 목적 함수로 wav2vec 2.0 오디오 임베딩에 대해 학습된 심층 네트워크를 통해 비침습적 뇌자도(MEG) 기록에서 검색할 수 있다. 그러나 해당 네트워크의 가중치는 전기생리학적 양과 직접적으로 대응되지 않으며, 어떤 음성 속성이 검색을 유도하는지도 여전히 불분명하다. 본 연구는 고성능 MEG-대-오디오 검색 아키텍처를 기반으로 하되 전단부와 디코더를 모두 재설계한다. 기존의 공간 주의 메커니즘은 평면화된 센서 배치에 작동하지만, 본 연구는 이를 3차원 MEG 헬멧 형상에 정의된 구면 조화 함수로 대체한다. 피험자 특이적 표현을 270개에서 25개 분기(branch)로 줄이고, 각 분기에 시간 필터를 추가하여 공간적·시간적으로 신경원 소스와 정합시키며, 합성곱 디코더는 더 얕게 구성한다. 훈련 전에 안구 및 심장 성분을 제거하여 자극 고정 지름길(shortcut)의 위험을 줄인다. MEG-MASC에서 이 모델은 1,005개 후보 중 39.75 ± 0.34%의 Top-1 정확도를 달성하며, 6개의 훈련된 솔루션에 걸쳐 약 20배 더 적은 디코더 파라미터를 사용한다. 모델 가중치는 소스 공간에 매핑되어 언어 지각 네트워크와 일치하는 생성원(generator)을 복구하며, 좌반구 편측화된 분기는 우반구에서는 뚜렷하지 않은 더 높은 주파수의 리듬 성분을 운반한다. 짝지어진 MEG 가림(occlusion) 실험에서는 19개 자극 특징 중 15개가 검색에 기여하며, 침묵, 소리 강도, 모음, 음향 시작점에서 가장 큰 효과가 나타난다. 무작위 단어 목록은 반대 양상을 보인다. 즉, 내러티브 MEG를 무작위 단어 목록에 대입하면 검색이 향상되는데, 이는 내러티브 구조가 없는 활동이 일관된 음성 중의 활동보다 복구 가능한 정보를 덜 담고 있음을 시사한다. wav2vec 표적은 정확도 손실 없이 약 12개의 학습된 특징 차원으로 축소될 수 있는 반면, 강한 시간 압축은 명확한 손실을 초래한다. 종합하면, 소스 매핑과 입력 개입은 검색을 유도하는 요인이 무엇인지를 밝혀준다.
English
Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto electrophysiological quantities, and it remains unclear which speech properties drive retrieval. We build on a high-performing MEG-to-audio retrieval architecture but redesign both its front end and decoder. Its spatial attention operates on a flattened sensor layout; we replace it with spherical harmonics defined on the three-dimensional MEG helmet geometry. We reduce the subject-specific representation from 270 to 25 branches, add a temporal filter to each branch to match it to a neuronal source in space and time, and make the convolutional decoder shallower. Ocular and cardiac components are removed before training to reduce the risk of stimulus-locked shortcuts. On MEG-MASC, the model reaches 39.75 +/- 0.34% Top-1 accuracy among 1005 candidates across six trained solutions, with about 20 times fewer decoder parameters. Its weights map to source space, recovering generators consistent with the speech-perception network, while left-lateralized branches carry higher-frequency rhythmic components not evident on the right. Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets. Random word lists behave oppositely: substituting narrative MEG into them improves retrieval, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech. The wav2vec target can be reduced to about twelve learned feature dimensions without loss of accuracy, whereas strong temporal compression causes a clear loss. Together, source mapping and input interventions reveal what drives retrieval.