ChatPaper.aiChatPaper

체화된 조작을 위한 데이터 피라미드

Data Pyramid for Embodied Manipulation

July 27, 2026
저자: Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Duo Zheng, Tianxing Chen, Jiaming Liu, Ziang Cao, Yunfan Lou, Wei Chow, Xian Sun, Yingshuo Wang, Kuangzhi Ge, Xiaowei Chi, Xidong Zhang, Zhibo Pang, Yiwu Zhong, Sirui Han, Zhihe Lu, Weihao Yuan, Qifeng Chen, Michael Yu Wang, Yao Mu, Ziwei Liu, Jianfei Yang, Ping Luo, Shanghang Zhang
cs.AI

초록

멀티모달 기반 모델들은 인터넷 전체를 소비함으로써 보고 말하는 법을 학습했다. 그러나 신체화된 에이전트(embodied agents)에게는 그러한 지름길이 존재하지 않는다. 이는 관찰과 물리적 상태 및 행동을 결합한 데이터를 필요로 하기 때문이다. 이러한 신호는 다양한 데이터 소스에 의해 정도의 차이는 있지만 제공될 수 있다. 본 연구에서는 신체화된 데이터 생태계를 다섯 가지 상호 보완적 소스(실제 로봇 데이터, UMI 스타일 데이터, 자아 중심적 및 외부 중심적 데이터, 시뮬레이션 데이터, 일반 시각-언어 데이터)를 아우르는 '피라미드'로 조직화한다. 이 피라미드는 확장성과 로봇 정합성(robot alignment) 사이의 긴장을 중심으로 구성되며, 각 소스를 데이터 품질, 다양성, 재사용성, 물리적 충실도 측면에서 추가로 특성화한다. 이후 최근의 신체화된 기반 모델들을 데이터 레시피(data recipe)의 관점에서 분석하여, 사전 훈련 중 각 소스가 어떻게 선택, 정렬, 혼합되는지 살펴본다. 신체화된 두뇌 모델, 시각-언어-행동 모델, 세계-행동 모델 모두에 대해 데이터 구성과 인식, 추론, 계획, 행동 생성, 세계 예측 능력 간의 관계를 정립한다. 마지막으로 여섯 가지 공개 과제, 즉 대규모 촉각 데이터셋 구축, 실패 및 복구 데이터 수집, 확장 가능한 데이터 수집 파이프라인 개발, 신체화 간 행동 정렬, 손재주 조작을 위한 자아 중심적 데이터 활용, 로봇 학습을 위한 원칙적인 데이터 레시피 설계를 논의하며 마무리한다. 본 연구가 차세대 신체화된 시스템 설계의 기초를 마련하기를 기대한다.
English
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a "pyramid" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.