ChatPaper.aiChatPaper

ACE-Data-0:以人為中心的環境感應擷取作為具身資料引擎

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

July 30, 2026
作者: Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu
cs.AI

摘要

具身智慧面臨著根本性的資料瓶頸。模型必須捕捉第一人稱感知、全身動作、靈巧操作、物體狀態、聲音與觸覺如何隨著人類在時間中追求目標而共同演化。現有資料集將這些經驗分散在不同的視角、模態或空間尺度中,使得完整的感知-行動迴圈僅能被部分觀察。我們提出環境捕捉引擎(Ambient Capture Engine, ACE),這是一套以人為中心的資料引擎,能將真實居家環境轉化為空間校準、時間同步的錄製工作室。ACE 在兩個互補的尺度上運作:桌面尺度配置可解析手部與物體的操作,而房間尺度配置則捕捉在佈置完整的居家環境中的全身動作、移動與互動。ACE 記錄自我中心視角與多視角外部視角的影片、全身與關節化手部動作、物體幾何與六自由度軌跡、音訊及觸覺訊號,形成統一的多元感官資料流。利用 ACE,我們建構了 ACE-Data-0,包含 150 小時、1,700 萬幀影片,涵蓋 200 個任務類別,由 50 位參與者在 2 個環境中執行,總共產生 75,000 個互動片段。該資料集涵蓋原子操作、長時程家務活動鏈,以及人與場景的互動,同時透過目標層級而非逐步指令來保留自然的行為變異。我們進一步提出一個層級式基準測試,從訊號逐步推進至場景要素,再進展到互動。對最新方法的評估揭示了在接觸、遮擋、自我運動及長時程時間範圍下仍存在重大差距。ACE-Data-0 提供帶有對齊的感知、運動學與接觸監督的同步人類示範,為模仿學習、世界模型、視覺-語言-行動系統及具身智慧提供了可擴展的基礎。
English
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.