OmniTacTune: 시각 정책의 촉각 잔차 적응을 위한 정책 비의존적 실세계 강화학습
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
July 4, 2026
저자: Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han, Ruohan Gao
cs.AI
초록
인간 영상, 원격 조작 및 로봇 시연으로부터 학습된 시각적 정책은 확장 가능한 운동 사전을 제공하지만, 성공이 국소 힘과 접촉 형상에 크게 의존하는 접촉이 많은 조작에서는 종종 실패합니다. 촉각 감지는 이러한 상보적 신호를 제공하지만, 촉각 데이터는 수집 비용이 많이 들고 센서, 로봇, 작업 간에 일반화하기 어렵습니다. 본 논문에서는 정책에 구애받지 않는 실제 환경 강화 학습 파이프라인인 OmniTacTune을 소개합니다. 이는 잔차 보정을 통해 사전 학습된 시각적 정책에 촉각 피드백을 적응시킵니다. OmniTacTune은 2단계 설계를 사용합니다. 먼저 자율 기본 정책 롤아웃으로부터 촉각 인식 학습을 부트스트래핑한 후, 온라인 상호작용을 통해 경량의 촉각 잔차 정책을 학습합니다. 광범위한 실험을 통해 OmniTacTune이 다양한 접촉이 많은 작업, 시각적 기본 정책 및 촉각 표현에 걸쳐 일반화됨을 보여줍니다. 네 가지 실제 환경 접촉이 많은 작업에서 40-80분 이내에 시각적 기본 정책의 성공률을 5-40%에서 85-100%로 향상시켜, 확장 가능한 시각적 로봇 정책에 촉각 피드백을 적응시키는 효율적인 경로를 입증합니다. 프로젝트 페이지: https://colinyu1.github.io/omnitactune-site/
English
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry. Tactile sensing provides these complementary signals, yet tactile data remain costly to collect and hard to generalize across sensors, robots, and tasks. We introduce OmniTacTune, a policy-agnostic real-world RL pipeline that adapts tactile feedback to pretrained visual policies through residual correction. OmniTacTune uses a two-stage design: it first bootstraps tactile-aware learning from autonomous base-policy rollouts, then learns a lightweight tactile residual policy through online interaction. Extensive experiments show that OmniTacTune generalizes across diverse contact-rich tasks, visual base policies, and tactile representations. Across four real-world contact-rich tasks, it improves visual base policies from 5-40% success to 85-100% within 40-80 minutes, demonstrating an efficient path for adapting tactile feedback to scalable visual robot policies. Project page: https://colinyu1.github.io/omnitactune-site/