픽셀을 넘어: 비디오 사전 지식에서 4D 세계로
Beyond Pixels: From Video Priors to 4D Worlds
August 11, 2026
저자: Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang
cs.AI
초록
4D 생성은 텍스트나 이미지와 같은 조건으로부터 동적 3D 장면을 합성한다. 기존 방법들은 생성된 RGB 비디오를 별도의 4D 모델로 재구성하거나, 특정 비디오 생성기를 변형하여 기하 구조를 직접 예측한다. 전자는 분포 불일치와 오류 전파 문제를 겪는 반면, 후자는 4D 예측을 특정 생성기에 종속시키며 생성기나 조건 설정 방식이 변경될 때 재훈련이 필요할 수 있다. 본 연구는 변분 오토인코더(VAE)를 공유하는 비디오 모델들의 최종 노이즈 제거 잠재 변수가 명시적 4D 예측을 위한 재사용 가능한 인터페이스를 제공할 수 있는지 묻는다. 이러한 통찰을 바탕으로, 본 연구는 RGB를 우회하여 비디오 잠재 변수를 사전 훈련된 4D 디코더의 토큰 그리드에 정렬하고 프레임별 및 전역 시공간 어텐션을 통해 이를 정제하는 직접 잠재-대-4D 생성을 도입하고, 이를 Latent-to-4D로 구현한다. 약 1K개의 기존 재구성 클립으로 훈련된 단일 체크포인트는 동일한 VAE 계열 내의 여러 비디오 확산 트랜스포머에 걸쳐 변경 없이 전이된다. Text4D-200 및 I4D-200에서 Latent-to-4D는 동일 잠재 변수를 사용하는 Wan+4RC 캐스케이드를 프로젝션 기반 DINO-F1에서 각각 2.88–3.45점 및 5.81점 능가하며, 기하 구조, 시간적 안정성, 전반적 품질 측면에서 인간 평가자들의 선호도 또한 더 높았다.
English
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.