N_0-VTLA:基于潜在触觉令牌的视觉-触觉-语言-动作模型扩展
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
July 26, 2026
作者: NeoteAI Team, Fudan TEAI Team
cs.AI
摘要
我们提出N_0-VTLA,一个视觉-触觉-语言-动作(VTLA)基础模型,能够实现:(1)具备触觉感知与触觉反馈控制的细粒度接触丰富操作;(2)利用已部署数据进行离线策略改进。在现有基于视觉的主干网络基础上,我们提出了一种触觉集成训练方案,包括视觉-触觉预训练、分阶段触觉通路集成和基于优势条件的离线策略改进。在预训练阶段,策略从NeoData(我们的大规模视觉-触觉机器人数据集)中学习广泛的接触先验;据我们所知,N_0-VTLA是首个在大规模触觉数据上预训练的VTLA模型。在后训练阶段,我们为策略添加一条预测性触觉通路,将在规模数据上学到的接触模式提炼为下游以触觉为中心的操作所需的精细动作调整。对于离线策略改进,我们提出了ALTER,一种基于优势条件的离线强化学习方法,它将相对进度与轨迹事件对比转换为二值优势标签,用于在固定部署语料库上进行策略训练,从而进一步提升接触丰富技能(如可变形物体操作)中的特定任务学习。在接触丰富基准测试中,N_0-VTLA以显著优势超越强基线:它赢得了NeoReal全部九项真实机器人任务,并在包含二十项任务的仿真套件上达到63.8%的平均成功率,而最强基线仅为44.0%。使用ALTER训练的N_0-VTLA策略在三个长时程真实机器人任务上达到75%至95%的成功率。这些结果为通用的触觉驱动操作策略奠定了基础。
English
We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, N_0-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, N_0-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. N_0-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.