OmniVAE: 공동 생성을 위한 교차 모달 정렬 기반 오디오-비디오 VAE
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
July 26, 2026
저자: Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
cs.AI
초록
최근 생성 모델들은 무음 비디오나 독립형 오디오 합성을 넘어, 동기화된 오디오와 비디오의 공동 생성을 지향하고 있다. 이러한 발전에도 불구하고, 정밀한 교차 모달 대응을 갖춘 오디오와 비디오의 공동 생성은 두 양식의 근본적인 구조적 차이로 인해 여전히 어려운 과제이다. 대부분의 기존 방법들은 별도로 학습된 오디오 및 비디오 VAE를 사용한다. 그 결과, 두 잠재 공간은 교차 모달 정렬이 부족하여, 다운스트림 생성 모델이 처음부터 교차 모달 동기화를 학습해야 한다. 우리는 오디오와 비디오 잠재 표현 간의 정밀한 의미 정렬을 학습하는 공동 학습 오디오-비디오 VAE인 OmniVAE를 제안한다. OmniVAE는 재구성 외에도, 시간-의미 대응을 포착하고 두 잠재 공간을 정렬하기 위해 세그먼트 수준의 오디오-비디오 대조 목표를 사용한다. 이와 동시에, 사전 학습된 양식별 의미 인코더의 특징을 각 양식에 증류하여 두 잠재 공간의 다운스트림 학습 가능성을 향상시킨다. 광범위한 실험을 통해 두 목표 모두 잠재 공간의 학습 가능성을 일관되게 향상시키며, 이는 다운스트림 텍스트-오디오-비디오 생성에서 더 높은 생성 품질과 더 정확한 교차 모달 동기화로 이어짐을 보여준다. 이러한 발견은 통합 표현 학습이 전모달 모델링의 기초로서 중요함을 강조한다.
English
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1