DreamX-Creator: 2K 해상도 네이티브 오디오-비디오 생성의 대중화

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

August 31, 2026
저자: Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu
cs.AI

초록

최근 비디오 생성기는 오디오를 생략하거나 별도의 단계에서 합성하는 경우가 많아, 시각적 역학과 음향 사건의 상호 모델링에 제한이 있다. 본 논문에서는 7B 생성기를 중심으로 한 소형의 네이티브 통합 오디오-비디오 생성 시스템인 DreamX-Creator 1.0을 제시한다. 첫 번째 프레임과 텍스트 프롬프트를 조건으로 하여, 생성기는 양식별 오디오 및 비디오 스트림을 공동으로 노이즈 제거한다. 이 스트림들은 네트워크 전반부에서 독립적으로 처리되며, 후반부에서는 게이티드 크로스모달 어텐션(Gated Cross-Modal Attention)을 통해 결합되는데, 이때 토큰별 및 헤드별 출력 게이트가 각 활성 크로스모달 어텐션 헤드 출력을 조정한다. 통합 오디오-비디오 데이터 시스템은 시간적으로 정합적인 클립을 구축하고 필터링하며, 구조화된 멀티모달 주석을 생성하고, 클립을 기능 중심 데이터 풀로 구성한다. 점진적 공동 학습(Progressive Joint Training)은 두 단계의 오디오-비디오 사전 학습과 고품질 미세 조정(High-Quality Finetuning)으로 구성된다. 오디오-비디오 강화 학습(Audio-Video Reinforcement Learning)은 비디오, 오디오 및 크로스모달 피드백을 해당 스트림에 라우팅하는 양식 인지 멀티모달 피드백(Modality-Aware Multimodal Feedback)을 통해 생성기를 추가 학습시킨다. 고해상도 출력을 위해, 본 논문의 자기회귀 1-스텝 2K 정제 파이프라인(Autoregressive 1-Step 2K Refinement)은 양방향 다중 스텝 교사를 자기회귀 다중 스텝 정제기로 변환하고, 이를 시간적 청크당 단일 노이즈 제거 평가만을 요구하는 학생 모델로 증류한다. 전반적으로 DreamX-Creator 1.0은 최첨단 오픈소스 시스템과 경쟁력 있는 성능으로 네이티브하고 동기화된 오디오-비디오 생성을 달성한다. 본 논문은 소형 7B 생성기와 2K 정제기를 공개함으로써 네이티브 오디오-비디오 생성을 민주화하고, 통합 오디오-비디오 생성 모델링에 대한 향후 연구를 위한 접근 가능한 기반을 제공하고자 한다.
English
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
PDF905September 2, 2026