SwanTale: 指示タスクとゼロショットタスクのための統合マルチスピーカー音声・オーディオ生成
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
August 3, 2026
著者: Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin, Cheng Yang
cs.AI
要旨
音声・オーディオ生成は、アニメーション吹き替え、オーディオドラマ、映画、広告、ゲーム、ポッドキャスト、ショート動画制作などでしばしば必要とされる。このようなシナリオでは、クリエイターは参照録音なしで声をデザインし、自然言語で話者スタイルを制御し、環境や音響効果を備えた音響シーンをサポートし、後からデザインした声を再利用できることが求められる。したがって、インストラクションタスクとゼロショットタスクの両方において、マルチスピーカーの音声・オーディオ生成をサポートすることが重要である。インストラクションタスクでは、環境、話者スタイル、および詳細なコンテンツのキャプションが必要であり、一方ゼロショットタスクでは、参照音声と同一の詳細なコンテンツを使用する。我々はこれらのタスクにデータ面とモデル面の両方から取り組む。まず、生の音声・オーディオデータをクリーニングし、対象を絞った合成データを追加し、多様かつ正確な多層キャプションをアノテーションするSwanData-Captionを提案する。次に、ゼロショットタスクとインストラクションタスクの両方をサポートするマルチスピーカー高表現力音声・オーディオ生成モデルであるSwanTaleを提案する。高品質なマルチオーディオモダリティ生成をサポートするため、SwanVAEを導入する。さらに、報酬条件付き品質制御とEngram条件付け、そしてマルチタスク・マルチオーディオモダリティモデリングのためのUnified MoEを採用する。加えて、カリキュラム学習とGRPOポストトレーニングを用いて、モデルが段階的に学習し能力を強化できるようにする。実験結果は、SwanTaleが複数の主要なゼロショット指標とインストラクション指標でトップとなり、両タスクで最高の表現力スコアを達成し、マルチスピーカーの音声・オーディオを含む複雑なインストラクション生成をサポートすることを示している。デモは https://swanaigc.github.io/#swantale で閲覧できる。
English
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.