音声ブレイン・コンピュータ・インターフェースにおける共通のコミュニケーション指標
A Common Measure of Communication for Speech Brain-Computer Interfaces
September 2, 2026
著者: Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones
cs.AI
要旨
音声ブレイン・コンピュータ・インターフェース(音声BCI)は、神経活動を言語に変換することで、麻痺のある人々に発話の回復をもたらす道を提供し、より広くは、自然なヒューマン・コンピュータ・インタラクションの新たな形態を可能にする。しかし、この可能性にもかかわらず、この分野には共通の進歩の尺度が存在しない。なぜなら、システムによって使用するデータセット、記録方法、発話の種類、語彙が異なるため、報告されるスコアはほとんど比較できないからである。この測定問題の背後には、未解決の2つの問いがある:(i)音声BCIは、ユーザーが伝達できるようにすべき単語の分布はどのようなものか、(ii)システムはその分布からどれだけの情報を伝達できるか。本稿では、オープン語彙相互情報量(OVMI)を導出することにより、この2つに対処する。OVMIは、ユーザーが伝達したいと望む可能性のある単語に関する参照分布に対する、デコーダが伝達する情報量を測定する情報理論的量である。これにより、異なる語彙などの異なる条件下で測定された能力を、共通のコミュニケーション尺度で評価することができる。通常報告される精度、単語誤り率(WER)、およびその他の指標は、システムが対応している単語のみを対象に計算されるため、システムがユーザーの意図した発話のどれだけを伝達できるかを過大評価し得ることを示す。次に、OVMIを用いて既存システムを比較し、システムがユーザーの言語のどの程度をサポートするかと、それらの単語をどの程度正確にデコードするかとの間のトレードオフを明らかにする。さらに、こうした比較はユーザーが伝達することが期待される内容に依存すること、そしてOVMIを最大化する語彙の選択によって、3つの音声領域において精度が最大16.3%相対的に改善されることを示す。したがって、OVMIは音声BCIコミュニティに対し、異種のシステムを比較し、語彙設計を改善し、この分野の進歩を測定するための原理に基づいた方法を提供する。
English
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.