ChatPaper.aiChatPaper

OmniVAE:一种具有跨模态对齐的音频-视频VAE,用于联合生成

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

July 26, 2026
作者: Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
cs.AI

摘要

最近的生成模型正从无声视频或独立音频合成转向同步音视频的联合生成。尽管取得了进展,但由于音视频在根本结构上的差异,联合生成具有细粒度跨模态对应的音视频仍具挑战性。现有方法大多使用分别训练的音频和视频VAE,导致两个隐空间缺乏跨模态对齐,下游生成模型需从零学习跨模态同步。我们提出OmniVAE,一个联合训练的音频-视频VAE,学习音频和视频隐表示之间的细粒度语义对齐。除重建外,OmniVAE使用段级音视频对比目标来捕获时序-语义对应并对齐两个隐空间。同时,它从预训练的模态特定语义编码器中蒸馏特征到各模态,提升两个隐空间的下游可学习性。大量实验表明,两个目标均一致提升了隐空间的可学习性,进而转化为下游文本到音视频生成中更高的生成质量和更准确的跨模态同步。这些发现强调了学习统一表示作为全模态建模基础的重要性。
English
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1