ChatPaper.aiChatPaper

ACE-Data-0: 인간 중심 앰비언트 캡처의 체화된 데이터 엔진

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

July 30, 2026
저자: Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu
cs.AI

초록

체화된 지능은 근본적인 데이터 병목 현상에 직면해 있다. 모델은 인간이 시간이 지남에 따라 목표를 추구하면서 1인칭 지각, 전신 동작, 정교한 조작, 객체 상태, 소리, 촉각이 함께 진화하는 방식을 포착해야 한다. 기존 데이터셋은 이러한 경험을 시점, 모달리티, 공간적 규모에 따라 분할하여 전체 지각-행동 루프가 부분적으로만 관찰되게 한다. 우리는 실제 가정 환경을 공간적으로 보정되고 시간적으로 동기화된 녹화 스튜디오로 변환하는 인간 중심 데이터 엔진인 Ambient Capture Engine(ACE)을 소개한다. ACE는 두 가지 상호 보완적 스케일로 작동한다. 테이블 스케일 구성은 손-물체 조작을 정밀하게 포착하며, 룸 스케일 구성은 가구가 비치된 주택 전역에서 전신 동작, 이동, 상호작용을 포착한다. ACE는 1인칭 및 다중 시점 외부(exocentric) 비디오, 전신 및 관절 손 동작, 객체 기하 및 6자유도(6-DoF) 궤적, 오디오 및 촉각 신호를 통합 다중감각 스트림으로 기록한다. ACE를 사용하여 우리는 200개 작업 범주에 걸쳐 150시간 및 1,700만 개의 비디오 프레임으로 구성된 ACE-Data-0을 구축한다. 이 데이터셋은 2개 환경에서 50명의 참가자가 수행한 총 75,000개의 상호작용 에피소드를 포함한다. 데이터셋은 원자적 조작, 장기 지평선의 가사 활동 체인, 인간-장면 상호작용을 포괄하며, 단계별 지시가 아닌 목표 수준의 지시를 통해 자연스러운 행동 변이를 보존한다. 또한 우리는 신호에서 장면 구성 요소, 그리고 상호작용으로 진행되는 계층적 벤치마크를 도입한다. 최첨단 방법에 대한 평가는 접촉, 가림, 자기 운동(egomotion), 긴 시간적 지평선 조건에서 상당한 격차를 드러낸다. ACE-Data-0은 정렬된 지각, 운동학, 접촉 지도를 갖춘 동기화된 인간 시연을 제공하며, 모방 학습, 세계 모델, 비전-언어-행동 시스템, 체화된 AI를 위한 확장 가능한 기반을 제공한다.
English
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.