从世界到手腕:任务条件化的未来手腕建模用于细粒度机器人操作
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
August 5, 2026
作者: Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu
cs.AI
摘要
视觉-语言-动作(VLA)模型通常将主视角和腕部视角观测视为并行的视觉输入,忽略了它们在机器人操作中的不同作用。然而,精细操作受益于对腕部局部交互在全局任务上下文下如何演化的预期。为解决这一局限,我们提出了World-to-Wrist VLA(W2-VLA),一种面向精细机器人操作、具备任务条件化未来腕部建模能力的VLA模型。给定当前多视角观测和任务指令,W2-VLA将一组潜在建模令牌情境化为视觉-语言模型与腕部预测器之间的紧凑接口。在该接口及已观测腕部历史的条件下,预测器预测未来的腕部潜在表示,并将其转化为面向未来感知的上下文用于动作预测。此外,我们引入了W2-CoT,一种合成流水线,用于生成描述操作进展、物理过渡线索及腕部局部证据的结构化标注。这些标注提供了辅助监督,以塑造任务条件化的潜在接口。在LIBERO、RoboTwin 2.0及真实世界操作任务上的实验表明,该方法在单臂和双臂场景中均提升了精细操作和接触敏感操作性能,同时保持80 Hz以上的动作生成速率。
English
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.