Qwen-Music 技術報告
Qwen-Music Technical Report
July 13, 2026
著者: Jin Xu, Kangdi Wang, Ruibin Yuan, Shun Lei, Xiong Wang, Xize Cheng, Xueyao Zhang, Yang Zhang, Yiheng Chen, Yongqi Wang, Yue Wang, Zhifang Guo, Zihan Liu, Zijian Lin, Dake Guo, Hangrui Hu, Lei Xie, Linhan Ma, Wei Xue, Wenxiang Guo, Xinfa Zhu, Xipin Wei, Yangze Li, Yuanjun Lv, Yuxuan Wang, Yunfei Chu, Zhiyong Wu
cs.AI
要旨
本レポートでは、高い音楽性と忠実度を備え、完全なボーカル歌唱を含む楽曲を生成可能な強力な音楽生成モデル、Qwen-Musicを紹介します。Qwen-Musicは2つのコアタスクをサポートします。テキストから音楽を生成する「Text to Music Generation」は、テキスト記述、歌詞、音楽的属性から全く新しい楽曲を生成します。また、「Cover Song Generation」は、既存の楽曲を異なるスタイルやボーカル特性で再解釈します。アーキテクチャとして、Qwen-Musicは3つのコアコンポーネント(Qwen-Music-Tokenizer、Qwen-Music-LLM、Qwen-Music-Render)を統合しています。Qwen-Music-Tokenizerはオーディオを25Hzの単一コードブックストリームであるMusic Semantic Tokensに圧縮し、LLM予測のためのセマンティックおよびメロディ情報を保持します。これらのトークンに基づき、Qwen-Music-LLMは自己回帰的な音楽セマンティックモデリングを実行します。主要な新規性は、メロディトークンベースのChain-of-Thought(Melody-CoT)メカニズムであり、楽曲全体の生成前にメロディを計画することで、創造性、音楽性、構造的一貫性、および参照オーディオベースのメロディクローニングを向上させます。離散セマンティックトークンの忠実度の限界を克服するため、Qwen-Music-Renderは生成的ステレオレンダリングを実行し、音響的詳細を豊かにし、高忠実度のステレオ波形を生成します。最後に、Qwen-Music-LLMを数百の言語をカバーする500万時間以上の多言語音楽データで学習します。まず品質認識型の事前学習カリキュラムを適用し、次に段階的事後学習(教師あり初期化、オフラインDPO、オンラインGSPOから構成)を用いて、音楽性と指示追従能力をさらに向上させます。600の中国語と英語のプロンプトにおいて、Qwen-Musicは16の客観的な音楽性および音質評価指標のうち13で最先端の結果を達成しました。専門評価者も、主要なプロプライエタリシステムよりもQwen-Musicを好みます。カバーソング生成において、Qwen-Musicは主要なプロプライエタリシステムよりも正確に参照メロディを保持します。
English
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, Qwen-Music-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train Qwen-Music-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than leading proprietary systems.