N_0-VTLA: 潜在触覚トークンによる視覚-触覚-言語-行動モデルのスケーリング
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
July 26, 2026
著者: NeoteAI Team, Fudan TEAI Team
cs.AI
要旨
我々は、触覚知覚と触覚フィードバック制御による高精度な接触リッチ操作と、蓄積された展開データからのオフライン方策改善の両方を可能にする、視覚・触覚・言語・行動(VTLA)基盤モデルN_0-VTLAを提案する。既存の視覚ベースのバックボーンに基づき、視覚-触覚事前学習、段階的触覚経路統合、およびアドバンテージ条件付きオフライン方策改善から構成される触覚統合のための訓練手順を提案する。事前学習中に、方策は我々の大規模視覚-触覚ロボットデータセットであるNeoDataから広範な接触事前知識を学習する。我々の知る限り、N_0-VTLAは大規模な触覚データで事前学習された最初のVTLAモデルである。事後学習では、大規模に学習された接触パターンを、下流の触覚中心操作に必要な微細な動作調整へと蒸留する予測型触覚経路を方策に追加する。オフライン方策改善のために、固定展開コーパス上での方策学習に向けて、相対的な進捗と軌跡イベントの比較をバイナリアドバンテージラベルに変換するアドバンテージ条件付きオフライン強化学習手法であるALTERを導入する。これにより、変形可能物体操作などの接触リッチなスキルにおけるタスク固有の学習をさらに向上させる。接触リッチなベンチマーク群において、N_0-VTLAは強力なベースラインを大差で上回る。実ロボットNeoRealの9タスクすべてで勝利し、20タスクのシミュレーションスイートでは平均成功率63.8%を達成した。これは最強ベースラインの44.0%に対してである。ALTERを用いて訓練されたN_0-VTLA方策は、3つの長期的な実ロボットタスクで75〜95%の成功率に達する。これらの結果は、汎用的な触覚駆動操作方策の基盤を築くものである。
English
We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, N_0-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, N_0-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. N_0-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.