음성 뇌-컴퓨터 인터페이스를 위한 공통 의사소통 측정법
A Common Measure of Communication for Speech Brain-Computer Interfaces
September 2, 2026
저자: Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones
cs.AI
초록
음성 뇌-컴퓨터 인터페이스(speech BCI)는 신경 활동을 언어로 변환하여 마비 환자의 말하기 회복을 위한 경로를 제공하고, 보다 넓게는 자연스러운 인간-컴퓨터 상호작용의 새로운 형태를 가능하게 한다. 이러한 가능성에도 불구하고 이 분야에는 진전 정도를 공통으로 측정할 수 있는 지표가 없다. 시스템마다 사용하는 데이터셋, 기록 방법, 발화 유형, 어휘 집합이 서로 다르므로 보고된 점수들을 서로 비교할 수 있는 경우가 드물기 때문이다. 이러한 측정 문제의 근저에는 두 가지 미해결 질문이 있다: (i) 음성 BCI가 사용자가 의사소통할 수 있도록 지원해야 하는 단어 분포는 무엇이며, (ii) 이 분포에서 시스템이 전달할 수 있는 정보량은 얼마인가 하는 것이다. 본 연구는 사용자가 전달하고자 할 수 있는 단어들에 대한 기준 분포 하에서 디코더가 전달하는 정보량을 측정하는 정보이론적 척도인 개방 어휘 상호 정보량(OVMI)을 도출하여 두 질문에 모두 답한다. 이를 통해 어휘 집합이 다른 등 서로 다른 조건에서 측정된 성능을 공통된 의사소통 척도로 평가할 수 있다. 또한 통상 보고되는 정확도, 단어 오류율(WER), 그리고 시스템이 지원하는 단어에 대해서만 계산되는 다른 지표들은 사용자가 의도한 발화 중 시스템이 전달할 수 있는 정도를 과대평가할 수 있음을 보인다. 이어서 OVMI를 사용해 기존 시스템들을 비교하고, 시스템이 사용자가 구사하는 언어 중 어느 정도까지 지원하는지와 그 단어들을 얼마나 정확하게 디코딩하는지 사이의 상충 관계를 밝히며, 이러한 비교가 사용자가 전달할 것으로 기대되는 내용에 따라 달라진다는 것을 보여 준다. 아울러 OVMI를 최대화하도록 어휘를 선택하면 세 가지 음성 도메인에서 정확도가 최대 16.3% 상대적으로 향상된다는 점을 입증한다. 따라서 OVMI는 음성 BCI 연구계에 이질적인 시스템들을 비교하고, 어휘 설계를 개선하며, 분야의 진전을 측정할 수 있는 원리에 기반한 방법을 제공한다.
English
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.