ACE-Data-0:以人为中心的环境感知作为具身数据引擎
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
July 30, 2026
作者: Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu
cs.AI
摘要
具身智能面临根本性的数据瓶颈。模型必须捕捉第一人称感知、全身运动、灵巧操作、物体状态、声音与触觉如何随着人类在时间推移中追求目标而协同演化。现有数据集将这种经验割裂于不同视角、模态或空间尺度之间,使得完整的感知-行动环路仅被部分观测。我们提出环境感知采集引擎(ACE),一个以人为中心的数据引擎,可将真实家庭环境转变为空间校准、时间同步的录制工作室。ACE 在两种互补尺度上运行:桌面级配置解析手-物操作,而房间级配置则捕捉带家具家庭环境中的全身运动、位移行走及交互行为。ACE 将第一人称与多视角外视角视频、全身与关节化手部运动、物体几何与六自由度轨迹、音频及触觉信号记录为统一的多感官数据流。利用ACE,我们构建了ACE-Data-0数据集,包含150小时、1700万帧视频,覆盖200个任务类别,由50名参与者在2个环境中执行,共计75,000个交互片段。该数据集涵盖原子操作、家务活动的长时程链条以及人-场景交互,同时通过目标级而非逐步指令的方式保留自然的行为变异性。我们进一步引入一个层级化基准,从信号逐步推进到场景组件,再进阶到交互。对现有最优方法的评估揭示了在接触、遮挡、自我运动及长时间跨度下的显著差距。ACE-Data-0提供了具有对齐的感知、运动学与接触监督的同步人类演示,为模仿学习、世界模型、视觉-语言-动作系统及具身AI提供了可扩展的基础。
English
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.