ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu

arXiv:2607.28625 · 2026-07-31 공개 · arXiv · PDF

vision-language-action embodied-intelligence imitation-learning spatial-calibration long-temporal-horizon interaction-episodes hierarchical-benchmark multimodal-sensing

Abstract

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

한국어 요약

한 줄 요약

ACE-Data-0은 실내 환경에서 사람-객체 상호작용을 다중 센서로 동기화 기록한 대규모 데이터셋이다.

핵심 기여도

핵심 아이디어

ACE-Data-0은 인간의 일상적 상호작용을 **전면적, 다중 센서, 동기화된 방식**으로 기록함으로써, **Embodied AI 학습에 필요한 데이터 허브**를 구축하고자 한다. 기존 데이터셋은 **분산된 환경, 단일 모달, 짧은 시간**에 제한되어 있었으나, ACE는 **실제 주거 환경**에서 **테이블-스케일**과 **룸-스케일** 두 가지 설정을 통해 **자연스러운 행동**을 **동시 다중 센서 기록**으로 구현한다. 특히, **egocentric 영상과 exocentric 영상**을 **동기화**하고, **6-DoF 객체 궤적**과 **터치 신호**를 포함함으로써, **물리적 상태와 행동의 일관성**을 보장한다.

기술적 접근법

주요 결과

의의 및 한계

ACE-Data-0은 **Embodied AI 연구에 필수적인 데이터 기반**을 제공하며, **일상적 상호작용의 전체 퍼셉션-액션 루프**를 기록한 **최초의 데이터셋**이다. 특히, **동기화된 다중 센서 기록**과 **자연스러운 행동**을 통해, **이미테이션 학습, 월드 모델, 비전-언어-액션 시스템** 개발에 기여할 수 있다. 그러나, **데이터 수집 환경이 제한적**(2개 환경)이며, **더 복잡한 상호작용**(예: 다인 상호작용)은 **미포함** 상태이다. 또한, **자동 추적 기반 어노테이션**은 **수동 검증**보다 **정확도가 낮을 수 있음**.

실용적 활용

ACE-Data-0은 **로봇이 실제 주거 환경에서 물체를 조작하거나, 가정 내 태스크를 수행하는 시스템** 개발에 활용 가능하다. 예를 들어, **스마트 홈 기기 제어**, **로봇 조리**, **의료 보조 장치** 등에서 **실제 인간 행동을 기반으로 한 학습**이 가능하다. 또한, **비전-언어-액션 시스템** 개발에도 기초 자료로 사용될 수 있다.