言语脑机接口的通用通信度量
A Common Measure of Communication for Speech Brain-Computer Interfaces
September 2, 2026
作者: Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones
cs.AI
摘要
言语脑机接口(speech BCIs)将神经活动转化为语言,为瘫痪患者恢复言语能力提供了途径,并在更广泛的层面上实现了新型自然的人机交互。尽管前景广阔,该领域却缺乏统一的进展衡量指标,因为不同系统采用不同的数据集、记录方法、言语类型和词汇表,因此其报告的分数难以相互比较。这一度量问题背后存在两个尚未解决的关键问题:(i) 言语脑机接口应使用户能够交流何种词汇分布;(ii) 系统能够传递该分布中的多少信息。我们通过推导开放词汇互信息(OVMI)来同时解决这两个问题。OVMI是一种信息论量,用于衡量解码器相对于用户可能希望交流词汇的参考分布所传递的信息量。这使得在不同条件(如不同词汇表)下测得的能力能够在统一的交流尺度上进行评估。我们证明,通常报告的准确率、词错误率(WER)以及其他仅针对系统所支持词汇计算的指标,可能会高估系统所能传达的用户意图言语量。随后,我们使用OVMI对现有系统进行比较,揭示系统所支持的用户语言覆盖范围与其解码准确度之间的权衡,表明这些比较结果取决于用户预计要交流的内容,并证实选择使OVMI最大化的词汇表可在三个语音领域使准确率相对提升高达16.3%。因此,OVMI为言语脑机接口研究界提供了一种原理性的方法,用以比较异构系统、改进词汇表设计并衡量该领域的进展。
English
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.