SwanTale: 지시 기반 및 제로샷 태스크를 위한 통합 다중 화자 음성 및 오디오 생성
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
August 3, 2026
저자: Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang
cs.AI
초록
음성 및 오디오 생성은 애니메이션 더빙, 오디오 드라마, 영화, 광고, 게임, 팟캐스트, 숏폼 비디오 제작 등에서 자주 필요하다. 이러한 시나리오에서 창작자는 참조 녹음 없이 목소리를 설계하고, 자연어로 화자 스타일을 제어하며, 환경과 오디오 효과를 통해 음향 장면을 지원하고, 이후 설계된 목소리를 재사용해야 할 수 있다. 따라서 인스트럭트 및 제로샷 태스크 모두를 위한 다중 화자 음성 및 오디오 생성을 지원하는 것이 중요하다. 인스트럭트 태스크는 환경, 화자 스타일, 세밀한 콘텐츠에 대한 캡션을 필요로 하는 반면, 제로샷 태스크는 동일한 세밀한 콘텐츠와 함께 참조 오디오를 사용한다. 우리는 데이터와 모델 양측에서 이러한 태스크를 해결한다. 먼저, 원시 음성 및 오디오 데이터를 정제하고, 목표 지향적 합성 보강을 추가하며, 다양하고 정확한 다중 수준 캡션을 주석으로 제공하는 SwanData-Caption을 제안한다. 다음으로, 제로샷과 인스트럭트 태스크를 모두 지원하는 다중 화자 표현적 음성 및 오디오 생성 모델인 SwanTale을 제안한다. 고품질의 다중 오디오 모달리티 생성을 지원하기 위해 SwanVAE를 도입한다. 또한, 다중 태스크 및 다중 오디오 모달리티 모델링을 위한 통합 MoE와 함께 보상 조건부 품질 제어 및 Engram 조건화를 채택한다. 추가로, 커리큘럼 러닝과 GRPO 사후 훈련을 사용하여 모델이 점진적으로 학습하고 능력을 강화하도록 한다. 실험 결과는 SwanTale이 여러 주요 제로샷 및 인스트럭트 지표에서 선두를 차지하고, 두 태스크 모두에서 최고의 표현력 점수를 달성하며, 다중 화자 음성 및 오디오를 포함하는 복잡한 인스트럭트 생성을 지원함을 보여준다. 데모는 https://swanaigc.github.io/#swantale에서 확인할 수 있다.
English
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.