ChatPaper.aiChatPaper

Qwen-Music 기술 보고서

Qwen-Music Technical Report

July 13, 2026
저자: Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu
cs.AI

초록

본 보고서에서는 완전한 보컬 창법으로 고음질의 음악성 높은 노래를 생성할 수 있는 강력한 음악 생성 모델인 Qwen-Music을 소개합니다. Qwen-Music은 텍스트 설명, 가사 및 음악적 속성으로부터 완전히 새로운 노래를 창작하는 텍스트-음악 생성(Text to Music Generation)과 기존 노래를 다양한 스타일과 보컬 특성으로 재해석하는 커버곡 생성(Cover Song Generation)이라는 두 가지 핵심 작업을 지원합니다. 아키텍처 측면에서 Qwen-Music은 Qwen-Music-Tokenizer, Qwen-Music-LLM 및 Qwen-Music-Render라는 세 가지 핵심 구성 요소를 통합합니다. Qwen-Music-Tokenizer는 오디오를 25Hz 단일 코드북 스트림의 음악 의미 토큰(Music Semantic Tokens)으로 압축하여 LLM 예측을 위한 의미 및 멜로디 정보를 보존합니다. 이러한 토큰을 기반으로 Qwen-Music-LLM은 자기회귀적 음악 의미 모델링을 수행하며, 핵심 혁신은 전체 노래 생성 전에 멜로디를 계획하는 멜로디 토큰 기반 사고 사슬(Melody-CoT) 메커니즘으로, 창의성, 음악성, 구조적 일관성 및 참조 오디오 기반 멜로디 복제를 개선합니다. 이산적인 의미 토큰의 충실도 한계를 극복하기 위해 Qwen-Music-Render는 생성적 스테레오 렌더링을 수행하여 음향 세부 사항을 풍부하게 하고 고충실도 스테레오 파형을 생성합니다. 마지막으로, 수백 개의 언어를 포함하는 500만 시간 이상의 다국어 음악 데이터로 Qwen-Music-LLM을 학습시킵니다. 먼저 품질 인식 사전 학습 커리큘럼(quality-aware pre-training curriculum)을 적용한 후, 점진적 사후 학습(progressive post-training)을 통해 지도 초기화(supervised initialization), 오프라인 DPO(offline DPO), 온라인 GSPO(online GSPO)를 순차적으로 수행하여 음악성과 명령 수행 능력을 더욱 향상시킵니다. 600개의 중국어 및 영어 프롬프트에 대해 Qwen-Music은 16개 객관적 음악성 및 오디오 품질 지표 중 13개에서 최첨단 결과를 달성했습니다. 전문 평가자들 또한 주요 독점 시스템보다 Qwen-Music을 선호했습니다. 커버곡 생성의 경우, Qwen-Music은 주요 독점 시스템보다 참조 멜로디를 더 정확하게 보존합니다.
English
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.