WanSong v1.0 技术报告
WanSong v1.0 Technical Report
July 16, 2026
作者: Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou
cs.AI
摘要
音乐生成基础模型近期引起了业界的广泛关注。然而,在实现高效生成高保真长音频的同时支持可控性,仍是一项挑战。为应对这些需求,我们提出了WanSong——一种简单但强大的方法,用于生成长格式、商业级歌曲。与自回归模型和级联多阶段管线(例如自回归后接扩散)不同,WanSong是一种纯扩散模型,可直接生成时长可达5分钟的高保真多语言歌曲,并一次性输出两个音轨(人声和背景音乐)。此外,我们的扩散框架通过步骤蒸馏实现了更快的推理,并提供高效的微调和定制路径,以支持下游编辑任务。
English
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), WanSong is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.