ChatPaper.aiChatPaper

統一音頻智能而不退化文本智能

Unified Audio Intelligence Without Regressing on Text Intelligence

July 6, 2026
作者: Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang, Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping
cs.AI

摘要

音訊智慧涉及對音訊及語音的理解、推理與生成。在此研究中,我們推出了 Nemotron-Labs-Audex-30B-A3B(Audex),這是一個基於 Nemotron-Cascade-2-30B-A3B(一款強大的純文字 MoE 大型語言模型)所建構的統一音訊-文字大型語言模型。Audex 採用簡潔的統一設計,僅使用單一 Transformer 解碼器:音訊輸入經編碼後投射至文字嵌入空間,而生成過程中則將文字令牌與量化音訊輸出令牌視為同等處理。此架構實現了強效的音訊-文字融合、無縫的多模態生成,並能兼容標準的大語言模型訓練與推論基礎設施。在訓練方面,我們精心篩選了包含 1574 億個音訊令牌與 3205 億個文字令牌的音訊-文字資料集,並對這些資料進行多階段監督式訓練,隨後再進行純文字 Cascade 強化學習與多領域同軌蒸餾。Audex 在音訊理解、語音辨識與翻譯、文字轉語音、音訊生成及語音轉語音生成方面均達到當前最佳水準,同時在其純文字大語言模型骨幹的推理、對齊、知識、長上下文與代理能力上,僅有極少或完全沒有衰退。我們開放了模型檢查點,以促進開放式研究。
English
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.