ChatPaper.aiChatPaper

SwanTale:面向指令與零樣本任務的統一多說話者語音與音頻生成

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

August 3, 2026
作者: Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang
cs.AI

摘要

語音與音頻生成常用於動畫配音、有聲劇、電影、廣告、遊戲、播客與短影音製作等場景。在這些場景中,創作者可能需要在不依賴參考錄音的條件下設計聲音、以自然語言控制說話者風格、以環境與音效支援聲音場景,並在後續重複使用所設計的聲音。因此,同時支援多說話者語音與音頻生成,涵蓋指令任務與零樣本任務,具有重要意義。指令任務需要環境描述、說話者風格及細粒度內容的標註,而零樣本任務則使用參考音訊並搭配相同的細粒度內容。我們從資料與模型兩方面著手解決這些任務。首先,我們提出 SwanData-Caption,對原始語音與音頻資料進行清理、加入針對性的合成資料覆蓋,並標註多層級且多樣準確的描述。接著,我們提出 SwanTale,一個支援零樣本與指令任務的多說話者表現性語音與音頻生成模型。我們引入 SwanVAE 以支援高品質的多音頻模態生成。隨後,我們採用獎勵條件品質控制與 Engram 條件機制,搭配 Unified MoE 進行多任務與多音頻模態建模。此外,我們使用課程學習與 GRPO 後訓練,使模型逐步學習並強化其能力。實驗結果顯示,SwanTale 在多項關鍵零樣本與指令指標上領先,在兩項任務中均取得最佳表現性評分,並能支援涉及多說話者語音與音頻的複雜指令生成。示範影片可於 https://swanaigc.github.io/#swantale 參閱。
English
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.