ChatPaper.aiChatPaper

SoftVTBench:面向可变形物体操作的变形感知视觉-触觉数据集与基准

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

August 19, 2026
作者: Bowen Jing, Mingxin Wang, Ruiyang Hao, Chenchen Ge, Hanwen Shen, Junjie He, Yang Cui, Yiming Hou, Weitao Zhou, Jiawei Wang, Minglei Li, Dandan Zhang, Ding Zhao, Houde Liu, Xiaofan Li, Si Liu, Ping Luo, Haibao Yu
cs.AI

摘要

物理交互质量是可变形物体操作的核心,然而大多数基准仅评估任务是否成功。策略可能在完成任务的同时发生滑动或造成过度压缩。一个主要瓶颈是缺乏视觉-触觉数据集,这类数据集需在完整任务中将策略可见的接触观测与独立的物理真值配对。我们提出SoftVTBench,一个面向物理交互感知的可变形物体操作的视觉-触觉数据集,包含4,000条专家示范和50多个资源,涵盖体积可变形物体及视觉匹配的刚性孪生体。以20Hz频率,每个回合同步多视角RGB、双指触觉RGB和标记运动、本体感觉、语言、二值和连续夹爪动作,以及仅供评估器使用的有限元(FEM)状态。在此数据集基础上,我们建立了一个闭环基准,利用固定的物体特定校准来定义变形感知成功率(DSR),该指标仅当轨迹完成任务且峰值归一化变形保持在容差范围内时才判定为成功。在Diffusion Policy、π₀.₅和FastWAM中,所有12种分布内配置均包含违反变形容差的成功轨迹,占各配置成功轨迹数的0.7%至24%。在分布偏移下,视觉-触觉变体在所有六项策略—套件比较中取得更高的任务成功率,在五项中取得更高的DSR,而其分布内收益则优劣参半。这些结果表明,仅提供触觉并不能确保有效的多模态融合。因此,SoftVTBench提供了一个通用的视觉-触觉资源,不仅用于研究策略是否成功,还用于研究策略如何与可变形物体进行物理交互,以及触觉何时能改善这种交互。
English
Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or causing excessive compression. A primary bottleneck is the absence of visuo-tactile datasets that pair policy-visible contact observations with independent physical ground truth over complete tasks. We introduce SoftVTBench, a visuo-tactile dataset for physical-interaction-aware deformable-object manipulation. It contains 4,000 expert demonstrations and more than 50 assets, including volumetric deformable objects and visually matched rigid twins. At 20 Hz, each episode synchronizes multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, and binary and continuous gripper actions, alongside evaluator-only finite-element (FEM) states. Building upon this dataset, we establish a closed-loop benchmark that uses fixed object-specific calibration to define the Deformation-aware Success Rate (DSR), which counts a rollout as successful only when it completes the task and keeps peak normalized deformation within tolerance. Across Diffusion Policy, π_{0.5}, and FastWAM, all 12 in-distribution configurations contain successful rollouts that violate the deformation tolerance, accounting for 0.7--24% of each configuration's successes. Under distribution shift, visuo-tactile variants achieve higher task success in all six policy--suite comparisons and higher DSR in five, whereas their in-distribution benefits are mixed. These results show that making touch available does not by itself ensure effective multimodal fusion. SoftVTBench therefore provides a common visuo-tactile resource for studying not only whether a policy succeeds, but how it physically interacts with deformable objects and when touch improves that interaction.