DreamX-Creator:普及2K分辨率的原生音视频生成
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
August 31, 2026
作者: Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
cs.AI
摘要
近期視頻生成器往往省略音訊,或是在獨立階段合成音訊,限制了視覺動態與聲學事件之間的雙向建模。我們提出 DreamX-Creator 1.0,一套以 7B 生成器為核心的精簡原生聯合音訊-視頻生成系統。在以首幀及文字提示為條件的情況下,該生成器聯合去噪模態特化的音訊與視頻串流。這些串流在網路前半部分獨立處理,並在後半部分通過門控跨模態注意力進行耦合,其中以 token 級與注意力頭級輸出閘控調節每個啟用的跨模態注意力頭輸出。統一的音訊-視頻資料系統構建並過濾時間連貫的片段、產生結構化的多模態註解,並將片段組織為能力導向的資料池。漸進式聯合訓練包含兩個音訊-視頻預訓練階段,隨後進行高品質微調。音訊-視頻強化學習進一步以模態感知多模態回饋對生成器進行後訓練,將視頻、音訊及跨模態回饋分別導向對應的串流。為達成高解析度輸出,我們的自迴歸 1 步 2K 精煉流程將雙向多步教師模型改編為自迴歸多步精煉器,並將其蒸餾為每個時間片段僅需一次去噪評估的學生模型。整體而言,DreamX-Creator 1.0 實現了原生且同步的音訊-視頻生成,效能可與最先進的開源系統相抗衡。透過釋出我們精簡的 7B 生成器與 2K 精煉器,我們期望促進原生音訊-視頻生成的普及化,並為未來統一音訊-視頻生成建模研究提供可及的基礎。
English
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.