TacForcing:执行时触觉反馈的流式动作生成
TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
August 26, 2026
作者: Jianbo Zhou, Boyuan Zhao, Yuzheng Zhang, Yiyang Chen, Wenxin Chen, Qiuyue Li, Xiangyang Gu, Yuhan Cao, Xiao Xia, Yanzhe Hu, Zhijie Deng
cs.AI
摘要
接触丰富的操作需要适应在动作时间范围内可能发生显著变化的接触状态。然而,基于分块的视觉-语言-动作模型从执行前收集的观测中预测完整的动作块,导致触觉条件在执行过程中变得过时。现有的触觉反应方法通常依赖单独的高频控制器,这增加了架构和训练的复杂度。在本文中,我们提出了TacForcing,一种流式动作生成框架,能够有效地融合执行时的触觉反馈。TacForcing不采用单独的反应式控制器,而是用流式动作专家替代标准动作专家,根据执行过程中获取的不断演化的触觉观测来生成动作。TacForcing还引入了执行感知触觉注意力机制(EATA),将触觉条件限制在接近执行的动作上,从而减少触觉获取与动作执行之间的时间不匹配。在六个模拟UniVTAC任务和三个真实世界接触丰富操作任务中,TacForcing分别实现了65%和69%的平均成功率,在两种场景下均优于强基线方法。
English
Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reactive approaches typically rely on separate high-frequency controllers, which increase both architectural and training complexity. In this paper, we introduce TacForcing, a streaming action-generation framework that effectively incorporates execution-time tactile feedback. Instead of employing a separate reactive controller, TacForcing replaces the standard action expert with a streaming action expert to generate actions conditioned on the evolving tactile observations acquired during execution. TacForcing also introduces Execution-Aware Tactile Attention (EATA), which restricts tactile conditioning to actions nearing execution, thereby reducing the temporal mismatch between tactile acquisition and action execution. Across six simulated UniVTAC tasks and three real-world contact-rich manipulation tasks, TacForcing achieves average success rates of 65% and 69%, respectively, outperforming strong baselines in both settings.