ChatPaper.aiChatPaper

N_0-TWAM: 접촉이 많은 조작을 위한 촉각-네이티브 세계-행동 모델의 스케일링

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

July 26, 2026
저자: NeoteAI Team, Fudan TEAI Team
cs.AI

초록

우리는 접촉이 많은 조작을 위해 미래의 시각과 미래의 접촉을 모두 예측하는 촉각 기반 세계-행동 모델 N_0-TWAM을 제시한다. 우리가 아는 한, 이는 대규모로 훈련된 최초의 촉각 세계-행동 모델이며, 접촉이 많은 작업에서 강력한 성능을 보여준다. 우리는 6개의 로봇 플랫폼과 450개의 작업에 걸친 촉각이 풍부한 시연 데이터를 대상으로 시각-촉각 공동 훈련을 통해 N_0-TWAM을 대규모로 사전 훈련한다. 통합된 힘 기반 촉각 표현인 NeoForce를 사용하여 동작 생성을 조건화하는 물리적으로 근거한 접촉 신호를 형성한다. 장기간의 다단계 조작을 개선하기 위해, 작업 단계화를 위한 촉각 접촉 이벤트를 도입하고 실행 중에 이를 순차적으로 진행한다. 실시간 효율성을 위해, 비디오 예측을 위한 풀 폭 전문가와 후속 동작 및 촉각 예측을 위한 슬림 전문가를 결합한 비대칭 Mixture-of-Transformers 아키텍처를 채택한다. 실제 및 시뮬레이션 벤치마크에서의 평가는 다양한 접촉이 많은 작업에서 N_0-TWAM의 성능을 입증하며, 정밀한 촉각 및 동작 예측을 위한 데이터 확장의 이점을 보여준다. 요약하면, N_0-TWAM은 세계-행동 모델에 시각, 촉각, 동작을 예측하는 능력을 부여하여 개방형 접촉이 많은 작업에서 정밀한 조작을 위한 견고한 기반을 구축한다. 코드베이스와 모델 체크포인트는 촉각 기반 로봇 조작의 추가 연구와 개발을 촉진하기 위해 공개될 예정이다.
English
We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on contact-rich tasks. We pre-train N_0-TWAM at large scale with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments and 450 tasks. We use NeoForce, a unified force-based tactile representation, to form a physically grounded contact signal that conditions action generation. To improve long-horizon and multi-stage manipulation, we introduce tactile contact events for task staging and advance through them during execution. For real-time efficiency, we adopt an asymmetric Mixture-of-Transformers architecture that pairs a full-width expert for video prediction with slim experts for downstream action and tactile prediction. Evaluations on both real and simulated benchmarks justify the capabilities of N_0-TWAM across a range of contact-rich tasks, and demonstrate the benefit of data scaling for precise tactile and action prediction. In summary, N_0-TWAM endows a world-action model with predictive capabilities to foresee vision, touch and action, building a solid foundation for fine-grained manipulation on open contact-rich tasks. The codebase and model checkpoints will be made publicly available to foster further research and development in tactile-enabled robotic manipulation.