TacForcing:具執行時觸覺回饋的串流行動生成
TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
August 26, 2026
作者: Jianbo Zhou, Boyuan Zhao, Yuzheng Zhang, Yiyang Chen, Wenxin Chen, Qiuyue Li, Xiangyang Gu, Yuhan Cao, Xiao Xia, Yanzhe Hu, Zhijie Deng
cs.AI
摘要
接觸豐富的操作需要適應在動作時域內可能大幅演變的接觸狀態。然而,基於區塊的視覺-語言-動作模型會根據執行前收集的觀測結果預測完整的動作區塊,使得觸覺條件化在執行期間逐漸過時。現有的觸覺反應式方法通常依賴於獨立的高頻控制器,這增加了架構和訓練的複雜性。在本文中,我們提出TacForcing,一個流式動作生成框架,能有效整合執行期間的觸覺回饋。TacForcing並非採用獨立的反應式控制器,而是以流式動作專家取代標準動作專家,以執行期間獲得的即時觸覺觀測為條件來生成動作。TacForcing還引入了執行感知觸覺注意力(EATA),將觸覺條件化限制在接近執行的動作上,從而減少觸覺獲取與動作執行之間的時間錯配。在六個模擬UniVTAC任務和三個真實世界的接觸豐富操作任務中,TacForcing分別達到了65%和69%的平均成功率,在兩種設定下均優於強基線模型。
English
Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reactive approaches typically rely on separate high-frequency controllers, which increase both architectural and training complexity. In this paper, we introduce TacForcing, a streaming action-generation framework that effectively incorporates execution-time tactile feedback. Instead of employing a separate reactive controller, TacForcing replaces the standard action expert with a streaming action expert to generate actions conditioned on the evolving tactile observations acquired during execution. TacForcing also introduces Execution-Aware Tactile Attention (EATA), which restricts tactile conditioning to actions nearing execution, thereby reducing the temporal mismatch between tactile acquisition and action execution. Across six simulated UniVTAC tasks and three real-world contact-rich manipulation tasks, TacForcing achieves average success rates of 65% and 69%, respectively, outperforming strong baselines in both settings.