N_0-VTLA: 잠재 촉각 토큰을 활용한 비전-촉각-언어-행동 모델의 스케일링
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
July 26, 2026
저자: NeoteAI Team, Fudan TEAI Team
cs.AI
초록
우리는 촉각 지각 및 촉각 피드백 제어를 활용하는 미세하고 접촉이 많은 조작과 저장된 배포 데이터로부터의 오프라인 정책 개선이 가능한 시각-촉각-언어-행동(VTLA) 기반 모델인 N_0-VTLA를 제시한다. 현재의 비전 기반 백본을 기반으로, 우리는 시각-촉각 사전 학습, 단계적 촉각 경로 통합, 및 어드밴티지 조건부 오프라인 정책 개선으로 구성된 촉각 통합을 위한 학습 레시피를 제안한다. 사전 학습 동안 정책은 대규모 시각-촉각 로봇 데이터셋인 NeoData로부터 광범위한 접촉 사전 지식을 학습한다. 우리가 아는 한, N_0-VTLA는 대규모로 촉각 데이터에 사전 학습된 최초의 VTLA 모델이다. 후속 학습 동안 우리는 대규모로 학습된 접촉 패턴을 하위의 촉각 중심 조작에 필요한 미세한 운동 조정으로 증류하는 예측적 촉각 경로로 정책을 보강한다. 오프라인 정책 개선을 위해, 우리는 고정된 배포 코퍼스에서 정책 훈련을 위한 이진 어드밴티지 레이블로 상대적 진전 및 궤적-이벤트 비교를 변환하는 어드밴티지 조건부 오프라인 강화 학습 방법인 ALTER를 도입하여, 변형 가능한 객체 조작과 같은 접촉이 많은 기술에 대한 작업별 학습을 추가로 개선한다. 접촉이 많은 벤치마크 전반에 걸쳐 N_0-VTLA는 강력한 기준선을 큰 폭으로 능가한다. 즉, 9개의 실제 로봇 NeoReal 과제 모두에서 우승하고, 가장 강력한 기준선의 44.0%에 비해 20개 과제 시뮬레이션 스위트에서 평균 성공률 63.8%를 달성한다. ALTER로 훈련된 N_0-VTLA 정책은 세 가지 장기 실제 로봇 과제에서 75~95%의 성공률을 달성한다. 이러한 결과는 다재다능한 촉각 기반 조작 정책을 위한 기반을 마련한다.
English
We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, N_0-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, N_0-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. N_0-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.