ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu
arXiv:2607.28625 · 2026-07-31 공개 · arXiv · PDF
vision-language-action embodied-intelligence imitation-learning spatial-calibration long-temporal-horizon interaction-episodes hierarchical-benchmark multimodal-sensing
Abstract
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
한국어 요약
한 줄 요약
ACE-Data-0은 실내 환경에서 사람-객체 상호작용을 다중 센서로 동기화 기록한 대규모 데이터셋이다.
핵심 기여도
- ACE 시스템을 통해 실내 환경에서 **egocentric/exocentric 영상, 6-DoF 객체 궤적, 터치 신호** 등을 **동기화** 기록.
- **ACE-Data-0** 데이터셋은 150시간, 17M 프레임, 75,000개 상호작용 에피소드를 포함하며, 200개 태스크 범주를 커버.
- **3단계 계층적 벤치마크**를 제안: 신호 예측 → 신체/객체 복원 → 상호작용 추정.
- 기존 방법들이 **접촉, 가림, 장기적 시간** 조건에서 **성능 저하**를 드러냄.
핵심 아이디어
ACE-Data-0은 인간의 일상적 상호작용을 **전면적, 다중 센서, 동기화된 방식**으로 기록함으로써, **Embodied AI 학습에 필요한 데이터 허브**를 구축하고자 한다. 기존 데이터셋은 **분산된 환경, 단일 모달, 짧은 시간**에 제한되어 있었으나, ACE는 **실제 주거 환경**에서 **테이블-스케일**과 **룸-스케일** 두 가지 설정을 통해 **자연스러운 행동**을 **동시 다중 센서 기록**으로 구현한다. 특히, **egocentric 영상과 exocentric 영상**을 **동기화**하고, **6-DoF 객체 궤적**과 **터치 신호**를 포함함으로써, **물리적 상태와 행동의 일관성**을 보장한다.
기술적 접근법
- **ACE 시스템**은 **테이블-스케일**과 **룸-스케일** 두 가지 설정으로 구성됨.
- 테이블-스케일: **고해상도 터치 센서**와 **밀집된 카메라**로 **세부 손-객체 상호작용** 기록.
- 룸-스케일: **전체 환경을 덮는 센서**로 **전신 움직임, 이동, 객체 상호작용** 기록.
- **모든 신호**는 **광학 클록**을 통해 **밀리초 수준 정밀도로 동기화**.
- **마커-브리지드 캘리브레이션**으로 **정적/웨어러블 카메라**를 **공통 월드 프레임**에 등록.
- **ACE-Data-0**은 **50명 참여자**, **2개 환경**, **200개 태스크 범주**로 구성.
- **75,000개 상호작용 에피소드**, **17M 프레임**, **150시간 기록**.
- **자동 추적 기반의 어노테이션** 제공: **객체 메시, 6-DoF 포즈, 바운딩 박스, 운동 경로** 등.
주요 결과
- **ACE-Data-0**에서 기존 방법들을 평가한 결과, **접촉, 가림, 장기적 시간** 조건에서 **성능 저하**가 명확히 드러남.
- 예: **egocentric 영상 기반의 손 움직임 추정**에서 **가림** 상황에서 **정확도 감소**.
- **3단계 벤치마크**에서 **30개 이상의 대표적 방법** 평가.
- **터치 예측**: **비주얼 기반**에서 **터치 신호** 예측 실험.
- **신체/손 포즈 복원**: **추적 기반의 ground-truth**와 비교.
- **상호작용 추정**: **egocentric/exocentric 영상** 간 비교 가능.
- **전체적으로 기존 방법들이 실내 환경의 복잡성**에 대응하지 못함.
의의 및 한계
ACE-Data-0은 **Embodied AI 연구에 필수적인 데이터 기반**을 제공하며, **일상적 상호작용의 전체 퍼셉션-액션 루프**를 기록한 **최초의 데이터셋**이다. 특히, **동기화된 다중 센서 기록**과 **자연스러운 행동**을 통해, **이미테이션 학습, 월드 모델, 비전-언어-액션 시스템** 개발에 기여할 수 있다. 그러나, **데이터 수집 환경이 제한적**(2개 환경)이며, **더 복잡한 상호작용**(예: 다인 상호작용)은 **미포함** 상태이다. 또한, **자동 추적 기반 어노테이션**은 **수동 검증**보다 **정확도가 낮을 수 있음**.
실용적 활용
ACE-Data-0은 **로봇이 실제 주거 환경에서 물체를 조작하거나, 가정 내 태스크를 수행하는 시스템** 개발에 활용 가능하다. 예를 들어, **스마트 홈 기기 제어**, **로봇 조리**, **의료 보조 장치** 등에서 **실제 인간 행동을 기반으로 한 학습**이 가능하다. 또한, **비전-언어-액션 시스템** 개발에도 기초 자료로 사용될 수 있다.