ChatPaper.aiChatPaper

語意瓶頸:運用語意表徵進行非侵入式語音解碼

The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding

September 9, 2026
作者: Gilad D. Landau, Dulhan Jayalath, Oiwi Parker Jones
cs.AI

摘要

非侵入式語音解碼仍受神經記錄的低訊號雜訊比所限,這使得音素或個別詞彙的細粒度重建變得困難。受到神經科學證據啟發,即高階語意表徵分布於各皮質區域,並在較慢的時間尺度上演變,我們假設語意內容可能比低階聲學或詞彙特徵更適合作為非侵入式解碼的目標。我們提出 Brain2Semantics2Text,一種透過中介語意嵌入空間重建文本的方法。我們的模型將句子層級的 MEG 反應映射至語意流形,再將預測出的嵌入反轉為自然語言。此語意瓶頸使得無須詞層級對齊即可恢復高階意義。我們描述該方法的核心原理、其實作,以及用於緩解學習可靠的神經至語意映射之挑戰的策略。最後,我們與先前的非侵入式 Brain2Text 方法進行比較,並展示更佳的句子層級結果。
English
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.