從世界到手腕:任務條件化的未來手腕建模於細粒度機器人操作
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
August 5, 2026
作者: Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu
cs.AI
摘要
視覺-語言-行動(VLA)模型通常將主視角與手腕視角的觀察視為並行的視覺輸入,忽略了它們在機器人操作中的不同角色。然而,細粒度操作受益於預測手腕局部互動在全局任務情境下如何演變。為了解決此限制,我們提出了世界到手腕VLA(W2-VLA),這是一種用於細粒度機器人操作、具備任務條件化未來手腕建模的VLA模型。在給定當前多視角觀察與任務指令的情況下,W2-VLA將一組潛在建模令牌情境化,作為視覺-語言模型與手腕預測器之間的緊湊介面。在此介面與已觀察手腕歷史的條件下,預測器預測未來手腕潛在表徵,並將其轉化為具有未來感知的情境,用於行動預測。此外,我們引入了W2-CoT,這是一個合成流程,產生描述操作進展、物理轉變線索與手腕局部證據的結構化註解。這些註解提供輔助監督,用以塑造任務條件化的潛在介面。在LIBERO、RoboTwin 2.0及真實世界操作任務上的實驗證明,在單臂與雙臂設定中,細粒度與接觸敏感的操控表現皆有提升,同時維持高於80 Hz的行動生成速率。
English
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.