ChatPaper.aiChatPaper

Deform360: 변형 가능한 세계 모델을 위한 대규모 다중 시각-촉각 데이터셋

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

July 6, 2026
저자: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li
cs.AI

초록

물체 동역학 예측(즉, 세계 모델링)은 로봇 조작의 근본적인 과제이며, 변형 가능한 객체는 고차원 상태 공간과 복잡한 재료 특성으로 인해 특히 어려운 사례에 해당한다. 현재의 세계 모델은 2D 픽셀 공간 또는 보다 명시적인 3D 기하학적 공간에서의 동역학 학습이라는 두 가지 뚜렷한 패러다임을 통해 접근하고 있다. 그러나 이들 각각의 상대적 강점과 한계에 대한 체계적 이해는 다양하고 대규모의 실제 데이터가 부족하여 아직 불분명하다. 이를 해결하기 위해 우리는 Deform360을 제시한다. 이는 198가지 일상 물체, 1,980개의 상호작용 시퀀스, 그리고 41개의 전방위 카메라와 양손 촉각 그리퍼로부터 수집된 215시간 이상의 관찰 데이터를 포함하는 대규모 시각-촉각 데이터셋으로, 전역적 움직임과 접촉에 의한 국소적 변형을 모두 포착한다. 새로운 마커 없는 시각-촉각 3D 추적 파이프라인을 활용하여 조밀한 기하학과 움직임을 추출한 후, 2D 비디오 모델과 3D 입자 모델을 비교하며 현재 최첨단 세계 모델을 체계적으로 평가한다. 마지막으로, 변형 가능한 객체에 대한 로봇 계획 작업을 수행함으로써 데이터셋의 실제 적용 가능성을 보여주는 예비 시연을 제공한다. 우리의 분석은 구조적 사전 정보와 확장성 사이의 상충 관계에 대한 핵심 통찰을 밝혀내며, 일반화 가능한 변형 객체 중심 세계 모델링의 미래 연구를 위한 견고한 벤치마크를 제공한다. 프로젝트 웹사이트: https://deform360.lhy.xyz
English
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric space. A systematic understanding of their relative strengths and limitations remains elusive due to the lack of diverse, large-scale real-world data. To address this, we present Deform360, a large-scale visuotactile dataset featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Leveraging a novel markerless visuotactile 3D tracking pipeline to extract dense geometry and motion, we systematically evaluate current state-of-the-art world models, comparing 2D video models against 3D particle models. Finally, we provide a preliminary demonstration indicating the real-world applicability of our dataset by performing robot planning tasks on deformable objects. Our analysis reveals key insights into the trade-offs between structural priors and scalability, providing a solid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz