ChatPaper.aiChatPaper

N_0-TWAM: 接触を伴う操作のための触覚ネイティブ・ワールドアクションモデルのスケーリング

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

July 26, 2026
著者: NeoteAI Team, Fudan TEAI Team
cs.AI

要旨

我々は、接触リッチな操作のための触覚ネイティブな世界行動モデルN_0-TWAMを提案する。これは将来の視覚と将来の接触の両方を予測する。我々の知る限り、これは大規模に学習された最初の触覚世界行動モデルであり、接触リッチなタスクにおいて強い能力を示す。我々は、6つのエンボディメントと450のタスクにわたる触覚リッチなデモンストレーションを用いた視覚・触覚の共同トレーニングにより、N_0-TWAMを大規模に事前学習する。我々は、統一された力ベースの触覚表現であるNeoForceを用いて、行動生成を条件付ける物理的根拠に基づく接触信号を構築する。長期的かつ多段階の操作を改善するために、タスクを段階化する触覚接触イベントを導入し、実行中にそれらのイベントを順次進めていく。リアルタイム効率のために、ビデオ予測用のフル幅エキスパートと、下流の行動予測および触覚予測用のスリムなエキスパートを組み合わせた非対称なMixture-of-Transformersアーキテクチャを採用する。実機とシミュレーションの両方のベンチマークにおける評価は、多様な接触リッチタスクにわたるN_0-TWAMの能力を実証し、高精度な触覚予測と行動予測に対するデータスケーリングの利点を示す。要約すると、N_0-TWAMは、視覚・触覚・行動を予見する予測能力を世界行動モデルに与え、オープンな接触リッチタスクにおける微細な操作のための強固な基盤を構築する。コードベースとモデルチェックポイントは、触覚対応ロボット操作のさらなる研究と開発を促進するために公開される予定である。
English
We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on contact-rich tasks. We pre-train N_0-TWAM at large scale with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments and 450 tasks. We use NeoForce, a unified force-based tactile representation, to form a physically grounded contact signal that conditions action generation. To improve long-horizon and multi-stage manipulation, we introduce tactile contact events for task staging and advance through them during execution. For real-time efficiency, we adopt an asymmetric Mixture-of-Transformers architecture that pairs a full-width expert for video prediction with slim experts for downstream action and tactile prediction. Evaluations on both real and simulated benchmarks justify the capabilities of N_0-TWAM across a range of contact-rich tasks, and demonstrate the benefit of data scaling for precise tactile and action prediction. In summary, N_0-TWAM endows a world-action model with predictive capabilities to foresee vision, touch and action, building a solid foundation for fine-grained manipulation on open contact-rich tasks. The codebase and model checkpoints will be made publicly available to foster further research and development in tactile-enabled robotic manipulation.