ChatPaper.aiChatPaper

OmniVAE: 同時生成のためのクロスモーダルアラインメントを備えた音声・映像VAE

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

July 26, 2026
著者: Jun Zhan, Chen Yang, Yitian Gong, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
cs.AI

要旨

近年の生成モデルは、無音映像や単独の音声合成から、同期した音声と映像の同時生成へと進化している。この進展にもかかわらず、音声と映像の基本的な構造的差異から、微細なクロスモーダル対応を伴う同時生成は依然として困難である。既存の手法のほとんどは、音声用と映像用のVAEを別々に学習する。その結果、二つの潜在空間はクロスモーダルアライメントが不足し、ダウンストリームの生成モデルはゼロからクロスモーダル同期を学習せざるを得なくなる。我々は、音声と映像の潜在表現間の微細な意味的アライメントを学習する、共同学習された音声-映像VAEであるOmniVAEを提案する。再構成に加えて、OmniVAEはセグメントレベルの音声-映像対照的目的関数を用いて、時間的・意味的対応関係を捉え、二つの潜在空間を整列させる。これと並行して、事前学習されたモダリティ固有の意味エンコーダから各モダリティへ特徴を蒸留し、両方の潜在空間のダウンストリーム学習可能性を向上させる。広範な実験により、両目的関数が一貫して潜在空間の学習可能性を改善し、ダウンストリームのテキストから音声・映像への生成において、高い生成品質とより正確なクロスモーダル同期をもたらすことが示された。これらの知見は、オムニモーダルモデリングの基盤として統一表現を学習することの重要性を強調している。
English
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1