V-RAE: 생성을 위한 비디오 잠재 공간의 재고찰
V-RAE: Rethinking Video Latent Spaces for Generation
August 13, 2026
저자: Minghui Guo, Shengqiong Wu, Hao Fei
cs.AI
초록
잠재 비디오 생성은 오토인코더에 의존하여 생성 모델이 작동하는 간결한 공간을 정의한다. 비디오 오토인코더 아키텍처는 상당히 발전했지만, 그 잠재 공간은 여전히 주로 픽셀 수준 재구성에 최적화되어 있으며 고수준 의미론적 구조화를 제한적으로 제공한다. 그러나 재구성에 최적화된 잠재 공간이 반드시 생성 모델링에 적합한 것은 아니다. 우리는 고정된 비전 파운데이션 모델 표현 위에 간결한 생성 잠재 표현을 구축하는 비디오 표현 오토인코더 V-RAE를 제안한다. 가벼운 시간적 풀링 모듈은 의미 구조를 보존하면서 시간적 중복성을 제거하고, 비디오 디코더는 압축된 특징으로부터 연속적인 움직임을 재구성한다. 우리는 네 가지 대표적인 고정 인코더를 사용하여 비디오 재구성, 의미론적 프로빙, 클래스 조건부 생성에서 V-RAE를 평가한다. V-RAE는 K600에서 2.13 rFVD를 달성하여 평가된 모든 대규모 사전 학습 비디오 VAE를 능가한다. V-RAE의 잠재 표현은 기존 비디오 토크나이저의 잠재 표현보다 훨씬 더 많은 의미 정보를 유지한다. 일치된 생성 설정에서 우리의 최상의 변형은 UCF101과 K600에서 각각 117.86과 19.16의 gFVD 점수를 달성하면서 최대 6배 더 빠르게 수렴한다. 또한 우리는 재구성 품질만으로는 생성 유용성을 특성화하기에 불충분함을 보여주고, 하위 생성 품질과 더 신뢰성 있게 상관되는 시간적 일관성 진단 도구인 tFVD를 도입한다. 비디오 생성 외에도 V-RAE는 일치된 예측 설정에서 Wan 2.2 VAE 잠재 공간보다 Cityscapes에서의 미래 비디오 예측을 개선한다. 종합적으로, 실험은 고정된 의미론적 표현이 비디오 재구성, 생성 및 예측 모델링을 지원할 수 있음을 보여준다. 프로젝트 페이지: https://v-rae.github.io/.
English
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and provide limited high-level semantic organization. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations. A lightweight temporal pooling module removes temporal redundancy while preserving semantic structure, and a video decoder reconstructs continuous motion from the compressed features. We evaluate V-RAE with four representative frozen encoders on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, our best variant achieves gFVD scores of 117.86 and 19.16 on UCF101 and K600, respectively, while converging up to 6x faster}. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, the experiments show that frozen semantic representations can support video reconstruction, generation, and predictive modeling. The project page: https://v-rae.github.io/.