ChatPaper.aiChatPaper

VibeVoice-ASR-Streaming 기술 보고서

VibeVoice-ASR-Streaming Technical Report

September 2, 2026
저자: Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Jiajun Zhang, Xie Chen, Furu Wei
cs.AI

초록

기존의 화자 귀속 ASR(speaker-attributed ASR) 시스템은 음성 인식과 화자 분할을 별개의 두 작업으로 취급했습니다. 최근에는 VibeVoice-ASR과 같은 엔드투엔드 모델들이 이 두 작업을 단일 모델 안에서 통합했습니다. 그러나 기존 통합 모델들은 여전히 주로 오프라인 인식을 지원하여, 실시간 음성 비서와 에이전트의 저지연 요구를 충족하기 어렵습니다. 이 문제를 해결하기 위해 우리는 스트리밍 화자 귀속 ASR을 위한 최초의 LLM 기반 엔드투엔드 접근법 중 하나인 VibeVoice-ASR-Streaming을 제시합니다. 이 모델은 고정 크기의 오디오 청크, 소량의 룩어헤드(lookahead) 오디오, 이전 텍스트를 인터리빙합니다. 이를 통해 모델은 별도의 화자 분할 단계 없이도 음성이 도착하는 대로 “누가 무엇을 말했는지”를 생성할 수 있습니다. 전사 정확도 측면에서 우리의 7B 모델은 다섯 개의 평가 세트에서 가장 낮은 평균 WER/CER을 달성했습니다. 화자 귀속 측면에서는 13개 평가 설정 중 12개에서 최고 또는 공동 최고 성능을 달성했습니다. 또한 우리는 1.5B 및 7B 모델 가중치와 추론 코드를 공개합니다.
English
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ''who said what'' as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.