具身操作的數據金字塔
Data Pyramid for Embodied Manipulation
July 27, 2026
作者: Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Duo Zheng, Tianxing Chen, Jiaming Liu, Ziang Cao, Yunfan Lou, Wei Chow, Xian Sun, Yingshuo Wang, Kuangzhi Ge, Xiaowei Chi, Xidong Zhang, Zhibo Pang, Yiwu Zhong, Sirui Han, Zhihe Lu, Weihao Yuan, Qifeng Chen, Michael Yu Wang, Yao Mu, Ziwei Liu, Jianfei Yang, Ping Luo, Shanghang Zhang
cs.AI
摘要
多模態基礎模型通過消費整個網際網路學會了看與說。具身智能體則不容許此類捷徑,因為它們需要將觀測與物理狀態及動作相耦合的數據。這些信號可由多種數據源以不同程度提供。在本工作中,我們將具身數據生態系統組織為一個「金字塔」,涵蓋五種互補的數據源:真實機器人數據、UMI風格數據、自我中心與外部中心數據、模擬數據以及通用視覺-語言數據。我們圍繞可擴展性與機器人對齊之間的張力來組織此金字塔,並進一步從數據品質、多樣性、可重用性及物理保真度等維度刻畫每種數據源。接著,我們透過數據配方的視角分析近期具身基礎模型,探討在預訓練過程中如何選擇、對齊及混合不同數據源。無論是具身腦模型、視覺-語言-動作模型還是世界-動作模型,我們都將數據組成與感知、推理、規劃、動作生成及世界預測等方面的能力相關聯。最後,我們討論六個開放性挑戰:構建大規模觸覺數據集、收集失敗與恢復數據、開發可擴展的數據收集流程、對齊跨具身型的動作、利用自我中心數據進行靈巧操作,以及為機器人學習設計有原則的數據配方。希望本工作能為下一代具身系統的設計奠定基礎。
English
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.