ChatPaper.aiChatPaper

세계-손목: 정밀 로봇 조작을 위한 과제 조건부 미래 손목 모델링

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

August 5, 2026
저자: Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu
cs.AI

초록

비전-언어-행동(VLA) 모델은 종종 메인 뷰와 손목 뷰 관측값을 병렬적인 시각 입력으로 취급하며, 로봇 조작에서 두 관측값이 지니는 서로 다른 역할을 간과한다. 그러나 정밀 조작은 전역적 과제 맥락 하에서 손목 국소 상호작용이 어떻게 전개될지 예측함으로써 이점을 얻을 수 있다. 이러한 한계를 해결하기 위해, 본 논문에서는 과제 조건부 미래 손목 모델링을 적용한 정밀 로봇 조작용 VLA 모델인 W2-VLA(World-to-Wrist VLA)를 제시한다. 현재의 다중 뷰 관측값과 과제 지시가 주어졌을 때, W2-VLA는 비전-언어 모델과 손목 예측기 사이의 간결한 인터페이스 역할을 하는 일련의 잠재 모델링 토큰을 맥락화한다. 이 인터페이스와 관측된 손목 이력에 조건부로, 예측기는 미래 손목 잠재 표현을 예측하며, 예측된 표현은 행동 예측을 위한 미래 인식 컨텍스트로 변환된다. 또한, 본 연구는 조작 진행 상황, 물리적 전환 단서, 손목 국소 증거를 설명하는 구조화된 주석을 생성하는 합성 파이프라인인 W2-CoT를 도입한다. 이러한 주석은 과제 조건부 잠재 인터페이스를 형성하는 보조 감독을 제공한다. LIBERO, RoboTwin 2.0 및 실제 로봇 조작 과제에 대한 실험 결과, 단일 팔 및 양팔 설정 모두에서 정밀하고 접촉에 민감한 조작 성능이 개선되었으며, 행동 생성 속도는 80Hz 이상을 유지하였다.
English
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.