ACE-Data-0: 人間中心の環境的捕捉による具現化データエンジン
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
July 30, 2026
著者: Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu
cs.AI
要旨
身体化知能(Embodied Intelligence)は、根本的なデータボトルネックに直面している。モデルは、人間が時間をかけて目標を追求する過程で、一人称知覚、全身動作、巧緻操作、物体状態、音、触覚がどのように共進化するかを捉えなければならない。既存のデータセットは、こうした経験を視点、モダリティ、空間スケールの間で断片化しており、完全な知覚-行動ループは部分的にしか観測されていない。本稿では、実在の家庭環境を空間的に較正され、時間的に同期された録音スタジオへと変革する、人間中心のデータエンジンであるAmbient Capture Engine(ACE)を紹介する。ACEは、相互補完的な2つのスケールで動作する。テーブルスケール構成は手-物体間の操作を高解像度で捉え、ルームスケール構成は家具の備えられた家庭内での全身動作、移動、インタラクションを捉える。ACEは、一人称視点および多視点の三人称視点映像、全身および関節構造を持つ手の動き、物体の形状と6自由度軌道、音声、触覚信号を、統合されたマルチ感覚ストリームとして記録する。ACEを用いて、2つの環境において50名の参加者によって遂行された、200のタスクカテゴリにわたる150時間・1,700万映像フレーム、合計75,000のインタラクションエピソードから成るACE-Data-0を構築した。このデータセットは、原子的操作、長時間にわたる家事活動の連鎖、人間-シーン間インタラクションを網羅し、ステップごとの指示ではなく目標レベルの指示を通じて、自然な行動のばらつきを保持している。さらに、信号からシーン構成要素、そしてインタラクションへと段階的に進む階層的ベンチマークを導入する。最先端手法の評価により、接触、遮蔽、自己運動、長い時間的視野の下での実質的なギャップが明らかになった。ACE-Data-0は、整合された知覚的・運動学的・接触的教師信号を伴う同期された人間のデモンストレーションを提供し、模倣学習、ワールドモデル、視覚-言語-行動システム、および身体化AIのためのスケーラブルな基盤を提供する。
English
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.