ChatPaper.aiChatPaper

Puffin-World:ネイティブな3Dワールド状態を備えた統一マルチモーダルモデルのスケーリング

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

September 3, 2026
著者: Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy
cs.AI

要旨

我々は、外部オフラインモジュールに依存せずに、物理理解、空間シミュレーション、および3D世界の生成と再構築を統合する、統一的なマルチモーダルアーキテクチャであるPuffin-Worldを提案する。3D世界を確実に構築し、それと相互作用するために、我々のフレームワークは、多様なタスクと柔軟なモーションをサポートする統一的Omni-Camera表現とともに、物理(重力場と緯度)、幾何(深度)、外観(画像)という3つの固有の世界状態を共同でモデル化する。これらの状態のモデル化に加えて、将来フレームにわたって物理ダイナミクスを伝播させる戦略を導入する。絶対的なカメラ特性を実世界にグラウンディングすることにより、Puffin-Worldは物理的に一貫性があり視覚的に安定した世界生成を可能にする。さらに、単一の生成プロセス内で外観と幾何を結合し、各将来ビューの合成とその基盤となる幾何の再構築を共同で行う。この統合的パラダイムにより、模倣や自己校正型の世界探索など、複数のタスク間の相乗効果を必要とするインターリーブ型クローズドループアプリケーションが可能になる。Puffin-Worldを複雑なシナリオに拡張するため、多様かつ困難なモーションを特徴とする1500万の視覚・言語・カメラトリプレットと100万のトラジェクトリから構成されるPuffin-16Mを構築した。この分野のさらなる研究を促進するため、コード、モデル、およびデータセットを公開した。
English
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.