ChatPaper.aiChatPaper

Qwen-RobotManip Technisch Rapport: Afstemming ontgrendelt schaal voor fundamentmodellen van robotmanipulatie

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

June 17, 2026
Auteurs: Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, Jie Zhang, Jingyang Fan, Gengze Zhou, Qihang Peng, Chenxu Lv, Xiaoyue Chen, An Yang, Fei Huang, Junyang Lin, Dayiheng Liu, Jingren Zhou, Chenfei Wu, Xiong-Hui Chen
cs.AI

Samenvatting

Fundamentmodellen in taal en multimodaliteit bereiken sterke generalisatie door heterogene data te aligneren onder een uniforme formulering en op schaal te trainen. In dit rapport onderzoeken we of dit schaalrecept kan worden toegepast op robotmanipulatie om echte generalisatie te bereiken. Dit is uitdagend omdat manipulatiegegevens, in tegenstelling tot tekst, van nature heterogeen, duur om te verzamelen en smal in diversiteit zijn, wat het tegelijkertijd aligneren en schalen moeilijk maakt. We presenteren Qwen-RobotManip, een generaliseerbaar Visie-Taal-Actie fundamentmodel gebouwd op Qwen-VL. Qwen-RobotManip introduceert een uniform aligneringskader over de representatie-, bewegings- en gedragsdimensies van manipulatie, waardoor grootschalige multi-brontraining coherent wordt in plaats van conflicterend. Deze aligneringscapaciteit stelt Qwen-RobotManip op zijn beurt in staat om manipulatiegegevens te absorberen op een schaal die eerdere trainingsregimes niet konden ondersteunen. Een mens-naar-robot synthese-pijplijn zet egocentrische handdemonstraties om in robotbanen over 15 platforms, en een rigoureuze curatiepijplijn harmoniseert heterogene datasets. Met alleen opensource-datasets en menselijke video's zonder eigen dataverzameling construeert Qwen-RobotManip een ~38.100 uur durend pretrainingcorpus en vertoont het emergente generalisatiecapaciteiten, waaronder zero-shot instructieopvolging, robuustheid tegen verstoringen, reactief foutherstel en cross-embodiment transfer. We vinden dat standaard benchmarks er niet in slagen de pretrainingkwaliteit te vatten en in plaats daarvan OOD-instellingen aannemen, waaronder RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF en RoboTwin-XE. Qwen-RobotManip presteert aanzienlijk beter dan eerdere state-of-the-art modellen, waaronder π0.5, in alle OOD-instellingen, staat op de 1e plaats in RoboChallenge met een relatieve verbetering van 20%, en wordt gevalideerd op echte robotplatforms waaronder AgileX ALOHA, Franka, UR en ARX.
English
Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including π0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.