ChatPaper.aiChatPaper

World-to-Wrist:高精度ロボット操作のためのタスク条件付き将来の手首モデリング

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

August 5, 2026
著者: Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu
cs.AI

要旨

視覚-言語-行動(VLA)モデルは、主視点と手首視点の観測を並列的な視覚入力として扱うことが多く、ロボット操作におけるそれらの異なる役割を見落としている。しかしながら、高精度な操作は、グローバルなタスクコンテキストの下で手首の局所的インタラクションがどのように進展するかを予測することで恩恵を受ける。この限界に対処するため、我々はタスク条件付きの未来手首モデリングを備えたVLAモデルであるWorld-to-Wrist VLA(W2-VLA)を提案する。現在の多視点観測とタスク指示が与えられると、W2-VLAは潜在モデリングトークンの集合を視覚-言語モデルと手首予測器の間のコンパクトなインターフェースとして文脈化する。このインターフェースと観測された手首履歴に条件付けられて、予測器は未来の手首潜在表現を予測し、それが行動予測のための未来認識コンテキストへと変換される。さらに、操作の進展、物理的遷移の手がかり、および手首の局所的証拠を記述する構造化アノテーションを生成する合成パイプラインであるW2-CoTを導入する。これらのアノテーションは、タスク条件付き潜在インターフェースを形成する補助的な教師信号を提供する。LIBERO、RoboTwin 2.0、および実世界の操作タスクにおける実験により、単腕および両腕の両設定において高精度で接触に敏感な操作が改善されることが実証され、同時に80Hzを超える行動生成レートが維持される。
English
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.