Deform360: 変形可能なワールドモデルのための大規模マルチビュー視触覚データセット
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
July 6, 2026
著者: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li
cs.AI
要旨
物体の動力学の予測(すなわち世界モデリング)は、ロボット操作における基本的な課題であり、特に変形可能物体のモデリングは、その高次元な状態空間と複雑な材料特性のために極めて困難なケースとなっています。現在の世界モデルは、2次元ピクセル空間上でのダイナミクス学習か、より明示的な3次元幾何空間上でのダイナミクス学習という、2つの異なるパラダイムでこの問題に取り組んでいます。しかし、多様で大規模な実世界データが不足しているため、これらのパラダイムの相対的な強みと限界に関する体系的な理解は依然として困難です。
この問題に対処するために、我々はDeform360を提案します。これは、日常生活で使用される198個の物体、1,980のインタラクションシーケンス、41台の全方位カメラと両手触覚グリッパーによる215時間以上の観測データを含む大規模な視覚触覚データセットであり、全体的な動きと接触に起因する局所的な変形の両方を捉えます。新規のマーカーレス視覚触覚3Dトラッキングパイプラインを活用して高密度な幾何形状と動きを抽出し、現在の最先端世界モデルを体系的に評価し、2Dビデオモデルと3D粒子モデルを比較します。最後に、変形可能物体に対するロボット計画タスクを実行することで、我々のデータセットの実世界での適用可能性を示す予備的な実証を提供します。
我々の分析は、構造的先行知識とスケーラビリティの間のトレードオフに関する重要な洞察を明らかにし、一般化可能な変形可能物体中心の世界モデリングにおける将来の研究のための堅固なベンチマークを提供します。プロジェクトウェブサイト: https://deform360.lhy.xyz
English
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric space. A systematic understanding of their relative strengths and limitations remains elusive due to the lack of diverse, large-scale real-world data. To address this, we present Deform360, a large-scale visuotactile dataset featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Leveraging a novel markerless visuotactile 3D tracking pipeline to extract dense geometry and motion, we systematically evaluate current state-of-the-art world models, comparing 2D video models against 3D particle models. Finally, we provide a preliminary demonstration indicating the real-world applicability of our dataset by performing robot planning tasks on deformable objects. Our analysis reveals key insights into the trade-offs between structural priors and scalability, providing a solid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz