OmniVAE:跨模態對齊的音頻-視頻VAE用於聯合生成
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
July 26, 2026
作者: Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
cs.AI
摘要
近年來的生成模型已從無聲影片或獨立音訊合成,進展至同步生成音訊與影片的聯合生成。儘管如此,由於音訊與影片在基本結構上的差異,要生成具有細粒度跨模態對應的同步音訊與影片仍具挑戰性。現有多數方法使用各自獨立訓練的音訊與影片變分自編碼器(VAE),導致兩個潛在空間缺乏跨模態對齊,使得下游生成模型必須從頭學習跨模態同步。我們提出 OmniVAE,這是一個聯合訓練的音訊-影片 VAE,能夠學習音訊與影片潛在表徵之間的細粒度語義對應。除了重建之外,OmniVAE 採用片段層級的音訊-影片對比目標,以捕捉時間-語義對應並對齊兩個潛在空間。同時,它也從預訓練的模態特定語義編碼器中蒸餾特徵至每個模態,從而改善兩個潛在空間的下游可學習性。大量實驗顯示,這兩個目標一致地提升了潛在空間的可學習性,進而在下游文字到音訊-影片生成中轉化為更高的生成品質與更準確的跨模態同步。這些發現凸顯了學習統一表徵作為全模態建模基礎的重要性。
English
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1