具身操控的数据金字塔
Data Pyramid for Embodied Manipulation
July 27, 2026
作者: Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Duo Zheng, Tianxing Chen, Jiaming Liu, Ziang Cao, Yunfan Lou, Wei Chow, Xian Sun, Yingshuo Wang, Kuangzhi Ge, Xiaowei Chi, Xidong Zhang, Zhibo Pang, Yiwu Zhong, Sirui Han, Zhihe Lu, Weihao Yuan, Qifeng Chen, Michael Yu Wang, Yao Mu, Ziwei Liu, Jianfei Yang, Ping Luo, Shanghang Zhang
cs.AI
摘要
多模态基础模型通过摄取整个互联网数据学会了看和说。具身智能体则无法享有这种捷径,因为它们需要将观测与物理状态和动作相耦合的数据。多种数据源能以不同程度提供这类信号。在这项工作中,我们将具身数据生态系统组织成一个“金字塔”,涵盖五种互补来源:真实机器人数据、UMI式数据、第一人称和第三人称数据、仿真数据以及通用视觉-语言数据。我们围绕可扩展性与机器人对齐之间的张力构建这个金字塔,并进一步从数据质量、多样性、可复用性和物理保真度角度描述每种来源。随后,我们通过数据配方的视角分析近期具身基础模型,考察在预训练过程中如何选择、对齐和混合不同数据源。对于具身大脑模型、视觉-语言-动作模型以及世界-动作模型,我们将数据构成与感知、推理、规划、动作生成和世界预测等能力联系起来。最后,我们讨论六个未解决的挑战:构建大规模触觉数据集、收集失败与恢复数据、开发可扩展的数据收集管道、跨具身形态对齐动作、利用第一人称数据进行灵巧操作,以及为机器人学习设计原则性的数据配方。希望这项工作能为下一代具身系统的设计奠定基础。
English
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.