AuK 技術レポート:音声生成および編集のためのオープンソース基盤モデル
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
September 8, 2026
著者: Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen
cs.AI
要旨
我々はAuKを紹介する。これは、自然言語指示と音声コンテキストの共通インターフェースを通じて音声生成と編集を統合するオープンソースの基盤モデルである。この広範な機能セットをサポートするために、我々は約30.3億の指示-音声インスタンスと195万時間の実効的な教師信号を、5つのタスクファミリー(音声生成、内容編集、音声強調と分離、パラ言語編集、音響編集)にわたって構築する。AuKは、意味的条件付けのためのマルチモーダル大規模言語モデル、音響的条件付けのための音声・一般音響・音楽で共同学習されたVAE、そして生成のためにデュアルストリームMMDiTブロックに続いて統合されたシングルストリームDiTブロックを実行するハイブリッドRectified Flow Transformerを組み合わせる。訓練は生成のみのウォームアップから始まり、生成と編集の合同事前学習へと進む。次に、我々は補完的な事後学習戦略を適用する。オープンエンド編集のための人間フィードバック選好最適化と、音声生成のための報酬ベース強化学習である。推論コストを削減するために、我々はさらに一貫性初期化とタスクルーティングされたDecoupled DMDを用いてモデルを蒸留する。得られたAuK-Flashは、分類器フリーガイダンスなしで4ステップ推論を実行し、同一条件下でフルモデルに対して4.5倍の実時間高速化を達成する。実験は、ゼロショットおよび指示制御された音声生成と一般的な指示誘導編集において最先端の性能を実証し、信号レベルの復元タスクにおいても競争力を維持している。我々は再現性とさらなる研究を支援するために、ソースコードとモデル重みの両方を公開する。
English
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.