N_0-VTLA:利用潛在觸覺標記擴展視覺-觸覺-語言-行動模型
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
July 26, 2026
作者: NeoteAI Team, Fudan TEAI Team
cs.AI
摘要
我們提出 N_0-VTLA,這是一個視覺-觸覺-語言-動作(VTLA)基礎模型,能夠 (1) 具備觸覺感知與觸覺回饋控制的精細接觸豐富操作,以及 (2) 從已部署資料中進行離線策略改進。我們在現有的視覺骨幹基礎上,提出一套觸覺整合的訓練流程,包含視觸覺預訓練、分階段觸覺路徑整合,以及優勢條件化離線策略改進。在預訓練期間,策略從 NeoData(我們的大規模視觸覺機器人資料集)學習廣泛的接觸先驗;據我們所知,N_0-VTLA 是首個在大規模觸覺資料上預訓練的 VTLA 模型。在後期訓練期間,我們為策略增加一條預測性觸覺路徑,將在大規模資料中學習到的接觸模式,蒸餾為下游觸覺為中心操作所需的精細動作調整。在離線策略改進方面,我們引入 ALTER,一種優勢條件化離線強化學習方法,將相對進展與軌跡事件比較轉換為二元優勢標籤,用於在固定部署語料庫上進行策略訓練,進一步提升接觸豐富技能(如可變形物體操作)的任務特定學習。在接觸豐富的基準測試中,N_0-VTLA 以大幅差距優於強基線:它在所有九項真實機器人 NeoReal 任務中勝出,並在二十項任務的模擬套件中達到 63.8% 的平均成功率,而最強基線為 44.0%。使用 ALTER 訓練的 N_0-VTLA 策略在三項長時程真實機器人任務中達到 75-95% 的成功率。這些結果為多功能觸覺驅動的操作策略奠定了基礎。
English
We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, N_0-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, N_0-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. N_0-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.