ChatPaper.aiChatPaper

SwanTale:面向指令与零样本任务的统一多说话人语音与音频生成

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

August 3, 2026
作者: Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang
cs.AI

摘要

语音和音频生成在动画配音、有声剧、电影、广告、游戏、播客和短视频制作中有着广泛需求。在这些场景中,创作者可能需要在没有参考录音的情况下设计声音,通过自然语言控制说话人风格,支持带有环境和音效的声学场景,并在后续复用所设计的声音。因此,支持同时面向指令任务和零样本任务的多说话人语音与音频生成至关重要。指令任务需要对环境、说话人风格和细粒度内容进行描述,而零样本任务则使用参考音频以及相同的细粒度内容。我们从数据和模型两方面着手应对这些挑战。首先,我们提出SwanData-Caption,对原始语音和音频数据进行清洗,补充有针对性的合成数据覆盖,并标注多样且准确的多层级描述。接着,我们提出SwanTale,一个支持零样本和指令任务的多说话人表现力语音与音频生成模型。我们引入SwanVAE以支持高质量的多音频模态生成,并采用奖励条件质量控制和印记条件(Engram conditioning),结合统一的混合专家模型(Unified MoE)实现多任务和多音频模态建模。此外,我们利用课程学习和GRPO后训练,使模型逐步学习并强化自身能力。实验结果表明,SwanTale在多个关键零样本和指令指标上处于领先地位,在两项任务中均取得最佳表现力得分,并支持涉及多说话人语音和音频的复杂指令生成。演示可在https://swanaigc.github.io/\#swantale查看。
English
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.