ChatPaper.aiChatPaper

OmniTacTune: 策略無關的真實世界強化學習應用於視覺策略的觸覺殘差適應

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

July 4, 2026
作者: Kelin Yu, Haode Zhang, Harish Ravichandar, Yunhai Han, Ruohan Gao
cs.AI

摘要

從人類影片、遠端操作與機器人示範中學習的視覺策略,提供了可擴展的動作先驗,但在高接觸力的操作任務中時常失效,因為這類任務的成功與否高度取決於局部作用力與接觸幾何。觸覺感測能提供這些互補訊號,然而觸覺資料收集成本高昂,且難以在感測器、機器人與任務之間進行泛化。我們提出「OmniTacTune」,一種與策略無關的真實世界強化學習流程,透過殘差校正將觸覺回饋適應至預先訓練的視覺策略。OmniTacTune採用兩階段設計:首先從自主基礎策略的滾動探索中啟動觸覺感知學習,接著透過線上互動學習輕量化的觸覺殘差策略。廣泛的實驗顯示,OmniTacTune能泛化至多樣化的高接觸力任務、視覺基礎策略與觸覺表徵。在四項真實世界的高接觸力任務中,它將視覺基礎策略的成功率從5-40%提升至85-100%,且僅需40-80分鐘的訓練時間,展現了將觸覺回饋適應至可擴展視覺機器人策略的有效途徑。專案頁面:https://colinyu1.github.io/omnitactune-site/
English
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry. Tactile sensing provides these complementary signals, yet tactile data remain costly to collect and hard to generalize across sensors, robots, and tasks. We introduce OmniTacTune, a policy-agnostic real-world RL pipeline that adapts tactile feedback to pretrained visual policies through residual correction. OmniTacTune uses a two-stage design: it first bootstraps tactile-aware learning from autonomous base-policy rollouts, then learns a lightweight tactile residual policy through online interaction. Extensive experiments show that OmniTacTune generalizes across diverse contact-rich tasks, visual base policies, and tactile representations. Across four real-world contact-rich tasks, it improves visual base policies from 5-40% success to 85-100% within 40-80 minutes, demonstrating an efficient path for adapting tactile feedback to scalable visual robot policies. Project page: https://colinyu1.github.io/omnitactune-site/