ChatPaper.aiChatPaper

超越像素:從視訊先驗到4D世界

Beyond Pixels: From Video Priors to 4D Worlds

August 11, 2026
作者: Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang
cs.AI

摘要

4D生成從文字或影像等條件合成動態3D場景。現有方法要么使用單獨的4D模型重建生成的RGB影片,要么調整特定的影片生成器以直接預測幾何形狀。前者存在分布不匹配與誤差傳播的問題,而後者將4D預測綁定到特定生成器,當生成器或條件設定改變時可能需要重新訓練。我們探討能否將共享變分自編碼器(VAE)的影片模型之最終去噪潛變量,作為通往明確4D預測的可重用介面。基於此洞見,我們引入直接從潛變量到4D的生成方法,並將其實現為Latent-to-4D,該方法透過將影片潛變量與預訓練4D解碼器的token網格對齊,再以逐幀與全域時空注意力進行精煉,從而繞過RGB。僅在約1K段現有重建片段上訓練後,單一檢查點即可未經修改地遷移到同一VAE家族內的多個影片擴散Transformer。在Text4D-200和I4D-200基準上,Latent-to-4D在基於投影的DINO-F1指標上超越匹配的相同潛變量Wan+4RC級聯,分別高出2.88–3.45和5.81個百分點,同時在幾何形狀、時間穩定性和整體品質方面也更受人類評估者偏好。
English
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.