DreamX-Creator:普及2K分辨率的原生音视频生成

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

August 31, 2026
作者: Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
cs.AI

摘要

近年来的视频生成器往往省略音频,或在独立阶段合成音频,限制了对视觉动态与声学事件的相互建模。我们提出DreamX-Creator 1.0,一个以7B参数生成器为核心的轻量级原生联合音视频生成系统。该系统以首帧和文本提示为条件,对模态特化的音频流与视频流进行联合去噪。两条流在网络前半部分独立处理,在后半部分通过门控跨模态注意力进行耦合,其中令牌级和注意力头级输出门控分别调制每个活跃跨模态注意力头的输出。统一音视频数据系统构建并筛选时间连贯的视频片段,生成结构化的多模态标注,并将片段组织为面向能力的数据池。渐进式联合训练包含两个音视频预训练阶段,随后进行高质量微调。音视频强化学习进一步通过模态感知的多模态反馈对生成器进行后训练,将视频、音频及跨模态反馈分别路由至对应数据流。针对高分辨率输出,我们的自回归单步2K细化流水线将双向多步教师模型适配为自回归多步细化器,并将其蒸馏为每个时间块仅需一次去噪评估的学生模型。总体而言,DreamX-Creator 1.0实现了原生、同步的音视频生成,性能可与最先进的开源系统相媲美。通过发布我们轻量的7B生成器与2K细化器,我们致力于推动原生音视频生成的普及,并为统一的音视频生成建模未来研究提供可访问的基础。
English
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
PDF905September 2, 2026