ピクセルを超えて:ビデオ事前分布から4D世界へ
Beyond Pixels: From Video Priors to 4D Worlds
August 11, 2026
著者: Zihao Liu, Xiaolong Shen, Zhenglin Zhou, Ruijie Quan, Yi Yang
cs.AI
要旨
4D生成は、テキストや画像などの条件から動的な3Dシーンを合成する。既存手法は、生成されたRGBビデオを別の4Dモデルで再構成するか、特定のビデオ生成器を適応させて幾何形状を直接予測するかのいずれかである。前者は分布のミスマッチと誤差伝播の問題を抱える一方、後者は4D予測を特定の生成器に結び付け、生成器や条件付け方式が変わると再学習が必要になる可能性がある。我々は、変分オートエンコーダ(VAE)を共有するビデオモデルの最終的なデノイズ済み潜在表現が、明示的な4D予測への再利用可能なインターフェースを提供できるかどうかを問う。この洞察に基づき、我々は直接的な潜在表現から4Dへの生成を導入し、それをLatent-to-4Dとして具体化する。これはRGBを介さず、ビデオ潜在表現を事前学習済み4Dデコーダのトークングリッドと整列させ、フレーム単位およびグローバルな時空間アテンションによって精緻化する。約1Kの既存の再構成クリップで訓練された単一のチェックポイントは、同じVAEファミリー内の複数のビデオ拡散トランスフォーマーに変更なしで転移する。Text4D-200およびI4D-200において、Latent-to-4Dは、対応する同一潜在表現のWan+4RCカスケードを、投影ベースのDINO-F1でそれぞれ2.88~3.45ポイントおよび5.81ポイント上回り、また幾何形状、時間的安定性、全体的な品質の点で人間の評価者にも好まれる。
English
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video generator to predict geometry directly. The former suffers from distribution mismatch and error propagation, whereas the latter ties 4D prediction to a specific generator and may require retraining when the generator or conditioning regime changes. We ask whether the final denoised latents of video models that share a variational autoencoder (VAE) can instead provide a reusable interface to explicit 4D prediction. Building on this insight, we introduce direct latent-to-4D generation and instantiate it as Latent-to-4D, which bypasses RGB by aligning a video latent with the token grid of a pretrained 4D decoder and refining it through frame-wise and global spatiotemporal attention. Trained on roughly 1K existing reconstruction clips, a single checkpoint transfers unchanged across multiple video diffusion transformers within the same VAE family. On Text4D-200 and I4D-200, Latent-to-4D surpasses matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88--3.45 and 5.81 points, respectively, while also being preferred by human raters for geometry, temporal stability, and overall quality.