ChatPaper.aiChatPaper

AuK 기술 보고서: 음성 생성 및 편집을 위한 오픈소스 파운데이션 모델

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

September 8, 2026
저자: Ziyang Ma, Zhikang Niu, Wenming Tu, Tianrui Wang, Ruiqi Yan, Junxi Liu, Yanru Huo, Nickk Huang, Yang Liu, Qicong Xie, Zeyu Xie, Hui Wang, Haitao Li, Zixuan Jiang, Yalin Li, Jie Fang, Yifan Duan, Zeyue Tian, Guangzheng Li, Haina Zhu, Shuyi Wang, Jinwen Wang, Mingyu Cui, Tian Tan, Auden, Sen Liang, Steve Yves, Shan Yang, Liefeng Bo, Zilong Zheng, Kai Yu, Eng-Siong Chng, Xie Chen
cs.AI

초록

우리는 자연어 지시문과 오디오 컨텍스트의 공통 인터페이스를 통해 음성 생성과 편집을 통합하는 오픈소스 파운데이션 모델 AuK를 소개한다. 이러한 광범위한 기능 집합을 지원하기 위해, 우리는 음성 생성, 내용 편집, 향상 및 분리, 준언어적 편집, 음향적 편집의 다섯 가지 과제군에 걸쳐 약 30억 3천만 개의 지시문-오디오 인스턴스와 195만 시간의 효과적인 지도를 구축한다. AuK는 의미적 조건화를 위한 멀티모달 대규모 언어 모델, 음향적 조건화를 위해 음성, 일반 오디오, 음악에 공동으로 훈련된 VAE, 그리고 생성을 위해 이중 스트림 MMDiT 블록을 수행한 뒤 통합 단일 스트림 DiT 블록을 수행하는 하이브리드 rectified-flow Transformer를 결합한다. 훈련은 생성 전용 워밍업으로 시작하여 생성-편집 공동 사전 훈련으로 진행된다. 그런 다음 우리는 개방형 편집을 위한 인간 피드백 선호도 최적화와 음성 생성을 위한 보상 기반 강화 학습이라는 상호 보완적인 사후 훈련 전략을 적용한다. 추론 비용을 줄이기 위해, 우리는 일관성 초기화와 작업 라우팅된 Decoupled DMD로 모델을 추가로 증류한다. 그 결과로 얻어진 AuK-Flash는 classifier-free guidance 없이 4단계 추론을 수행하며, 동일 조건에서 전체 모델 대비 4.5배의 월클록 속도 향상을 달성한다. 실험은 제로샷 및 지시 제어 음성 생성과 일반 지시 기반 편집에서 선도적인 성능을 입증하는 한편, 신호 수준 복원 과제에서는 경쟁력 있는 성능을 유지한다. 우리는 재현성과 후속 연구를 지원하기 위해 소스 코드와 모델 가중치를 모두 공개한다.
English
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.