ChatPaper.aiChatPaper

Puffin-World: 네이티브 3D 세계 상태를 갖춘 통합 멀티모달 모델의 확장

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

September 3, 2026
저자: Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy
cs.AI

초록

우리는 물리 이해, 공간 시뮬레이션, 3D 세계 생성 및 재구성을 외부 오프라인 모듈에 의존하지 않고 통합하는 다중 모달 아키텍처인 Puffin-World를 제안한다. 3D 세계를 안정적으로 구축하고 상호작용하기 위해, 본 프레임워크는 다양한 작업과 유연한 움직임을 지원하는 통합 Omni-Camera 표현과 함께 물리(중력장 및 위도), 기하(깊이), 외관(이미지)이라는 세 가지 고유 세계 상태를 공동으로 모델링한다. 이러한 상태들의 모델링을 넘어, 우리는 미래 프레임들에 걸쳐 물리적 역학을 전파하는 전략을 도입한다. 절대 카메라 속성을 실제 세계에 근거 지음으로써 Puffin-World는 물리적으로 일관되고 시각적으로 안정적인 세계 생성을 가능하게 한다. 또한 우리는 단일 생성 과정에서 외관과 기하를 결합하여 각 미래 뷰를 공동으로 합성하고 그 기저 기하를 재구성한다. 이러한 통합 패러다임은 모방(mimic) 및 자가 교정(self-calibrated) 세계 탐험을 포함하여 여러 작업 간의 시너지를 요구하는 교차적 폐루프 응용을 가능하게 한다. 복잡한 시나리오로 Puffin-World를 확장하기 위해, 우리는 다양한 도전적 움직임을 포함하는 1,500만 개의 비전-언어-카메라 트리플릿과 100만 개의 궤적으로 구성된 Puffin-16M을 구축한다. 이 분야의 추가 연구를 촉진하기 위해 코드, 모델, 데이터셋을 공개한다.
English
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.