ChatPaper.aiChatPaper

Puffin-World:具有原生3D世界状态的统一多模态模型的扩展

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

September 3, 2026
作者: Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy
cs.AI

摘要

我们提出了Puffin-World,一种统一的多模态架构,它将物理理解、空间模拟以及3D世界生成与重建融为一体,且无需依赖外部离线模块。为了可靠地构建3D世界并与之交互,我们的框架联合建模三种原生世界状态:物理(重力场和纬度)、几何(深度)和外观(图像),并引入支持多样化任务和灵活运动的统一Omni-Camera表示。除了对这些状态进行建模外,我们还提出了一种跨未来帧传播物理动力学的策略。通过将绝对相机属性锚定于真实世界,Puffin-World实现了物理一致且视觉稳定的世界生成。我们进一步在单一生成过程中耦合外观与几何,联合合成每个未来视图并重建其底层几何。这一统一范式使得需要多任务协同的交错闭环应用成为可能,包括模仿(mimic)与自校准世界探索。为了将Puffin-World扩展到复杂场景,我们构建了Puffin-16M数据集,包含1500万个“视觉-语言-相机”三元组和100万条具有多样且富有挑战性运动的轨迹。为促进该领域的进一步研究,我们发布了代码、模型和数据集。
English
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.