身体性操作のためのデータピラミッド
Data Pyramid for Embodied Manipulation
July 27, 2026
著者: Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Duo Zheng, Tianxing Chen, Jiaming Liu, Ziang Cao, Yunfan Lou, Wei Chow, Xian Sun, Yingshuo Wang, Kuangzhi Ge, Xiaowei Chi, Xidong Zhang, Zhibo Pang, Yiwu Zhong, Sirui Han, Zhihe Lu, Weihao Yuan, Qifeng Chen, Michael Yu Wang, Yao Mu, Ziwei Liu, Jianfei Yang, Ping Luo, Shanghang Zhang
cs.AI
要旨
マルチモーダル基盤モデルは、インターネット全体を消費することで「見る」ことと「話す」ことを学習した。身体化エージェントにはそのような近道は存在しない。なぜなら、観測と物理状態および行動を結びつけるデータが必要だからである。これらの信号は、複数のデータソースによって様々な程度で提供されうる。本稿では、身体化データエコシステムを、5つの相補的なソース(実ロボットデータ、UMI形式データ、自己中心的および外部視点データ、シミュレーションデータ、汎用視覚言語データ)にわたる「ピラミッド」として整理する。このピラミッドを、スケーラビリティとロボット適応との間の緊張関係を中心に構成し、各ソースをデータ品質、多様性、再利用可能性、物理的忠実性の観点からさらに特徴づける。次に、近年の身体化基盤モデルをそのデータレシピの観点から分析し、事前学習中に異なるソースがどのように選択、調整、混合されるかを検討する。身体化脳モデル、視覚言語行動モデル、世界行動モデルのいずれにおいても、データ構成と知覚、推論、計画、行動生成、世界予測における能力との関連を考察する。最後に、6つの未解決課題、すなわち大規模触覚データセットの構築、失敗と回復データの収集、スケーラブルなデータ収集パイプラインの開発、身体化間での行動の調整、自己中心データを活用した巧みな操作、ロボット学習のための原理的なデータレシピの設計について議論する。本稿が次世代身体化システム設計の基盤を築くことを願っている。
English
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.