DreamX-Creator:2K解像度におけるネイティブ音声・動画生成の民主化
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
August 31, 2026
著者: Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
cs.AI
要旨
近年のビデオ生成器は、しばしば音声を省略するか、別の段階で合成することが多く、視覚的なダイナミクスと音響イベントの相互モデリングが制限されている。本論文では、7Bジェネレータを中核とする、コンパクトなネイティブ統合音声-ビデオ生成システムであるDreamX-Creator 1.0を提案する。初帧とテキストプロンプトを条件として、生成器はモダリティ特化型の音声ストリームとビデオストリームを共同でデノイジングする。これらのストリームは、ネットワークの前半では独立に処理され、後半ではゲート付きクロスモーダルアテンション(Gated Cross-Modal Attention)を通じて結合される。このアテンション機構では、トークン単位およびヘッド単位の出力ゲートが、各活性化クロスモーダルアテンションヘッドの出力を変調する。統合型音声-ビデオデータシステム(Audio-Video Data System)は、時間的に一貫性のあるクリップを構築・フィルタリングし、構造化されたマルチモーダルアノテーションを生成し、クリップを能力指向のデータプールに整理する。段階的統合訓練(Progressive Joint Training)は、2段階の音声-ビデオ事前学習と、それに続く高品質ファインチューニング(High-Quality Finetuning)で構成される。さらに、音声-ビデオ強化学習(Audio-Video Reinforcement Learning)により、モダリティ認識型マルチモーダルフィードバック(Modality-Aware Multimodal Feedback)を用いて生成器を事後訓練し、ビデオ関連、音声関連、およびクロスモーダルなフィードバックを対応するストリームにルーティングする。高解像度出力のため、我々の自己回帰型1ステップ2Kリファインメントパイプライン(Autoregressive 1-Step 2K Refinement pipeline)は、双方向マルチステップの教師モデルを自己回帰型マルチステップリファイナーへ適応させ、それを各時間チャンクにつき1回のデノイジング評価のみを必要とする学生モデルへ蒸留する。全体として、DreamX-Creator 1.0は、オープンソースの最先端システムと競合する性能を備えた、ネイティブで同期した音声-ビデオ生成を実現する。我々は、コンパクトな7Bジェネレータと2Kリファイナーを公開することで、ネイティブな音声-ビデオ生成の民主化を図り、統合音声-ビデオ生成モデリングにおける将来の研究のためのアクセス可能な基盤を提供する。
English
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.