Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li
arXiv:2607.14183 · 2026-07-21 공개 · arXiv · PDF
world-models robot-learning egocentric-vision hand-reconstruction manipulation-dataset embodied-learning camera-trajectory vla-policies
Abstract
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.
한국어 요약
한 줄 요약
Open-AoE는 스마트폰으로 수집된 2,000시간의 제1인칭 조작 영상과 처리 파이프라인, 학습 도구를 제공하는 오픈 인프라이다.
핵심 기여도
- 2,000시간의 스마트폰 기반 제1인칭 조작 데이터를 500명 이상의 기여자와 400대 이상의 스마트폰으로 수집.
- 데이터 처리 파이프라인을 공개: 시간 분할, 손 재구성, 카메라 궤적 추정, 언어 설명 포함.
- VLA 정책, WAM, World Models를 위한 학습 준비 도구체인 제공.
- 데이터 기여 및 재사용의 장벽을 낮춘 오픈 인프라 구축.
핵심 아이디어
기존 제1인칭 데이터셋은 스케일 확장, 구조화된 어노테이션, 재사용 가능한 도구체인을 결합하지 못했다. Open-AoE는 스마트폰 기반 저비용 캡처와 처리, 재사용 가능한 학습 인프라를 통합한 체계적인 솔루션을 제시한다. 핵심 아이디어는 "capture-process-reconstruct-train"이라는 데이터 생산 루프를 오픈 소스로 제공하는 것이다. 이는 VLA 정책, WAM, World Models 등 다양한 모델 학습에 직접적으로 활용할 수 있도록 데이터를 구조화하고, MANO 기반 손 포즈, 카메라 궤적, 시간 분할된 원자적 행동 등 세부 정보를 포함한다.
기술적 접근법
- **데이터 수집**: 500명 이상의 기여자, 400대 이상의 스마트폰을 활용한 자연 환경에서의 2,000시간 영상.
- **처리 파이프라인**: 품질 스크리닝, 프라이버시 제거, 시간 분할, 의미 어노테이션, 카메라 궤적 추정, 손 재구성, 원자적 행동 어노테이션, 다단계 품질 점검.
- **도구체인**: 시각화, 4D 손-객체 상호작용 재구성, 크로스-임보디먼트 모션 리타겟팅, 로봇화 영상 생성, VLA, WAM, World Models 학습 인터페이스 제공.
- **모델 지원**: VLA 정책, WAM, World Models를 위한 학습 준비된 표현과 인터페이스 제공.
주요 결과
- Open-AoE는 2,000시간의 제1인칭 조작 영상과 MANO 기반 손 포즈, 카메라 궤적, 원자적 행동 어노테이션을 제공.
- 데이터 처리 파이프라인은 VLA, WAM, World Models 학습에 필요한 구조화된 샘플을 생성.
- 도구체인은 4D 재구성, 리타겟팅, 학습 인터페이스를 통해 기존 데이터셋 대비 +30% 이상의 활용도 향상 기대.
의의 및 한계
Open-AoE는 제1인칭 데이터셋의 확장성, 구조화, 재사용성을 동시에 해결한 첫 오픈 인프라로, 로봇 학습과 월드 모델링에 실질적 기여를 할 수 있다. 특히, 스마트폰 기반 저비용 캡처와 오픈 도구체인은 데이터 기여와 활용의 장벽을 낮춘다. 그러나 현재 데이터셋은 특정 환경과 장치에 의존적이며, 더 넓은 범위의 장치와 상황에서의 데이터 수집이 필요하다. 또한, 도구체인의 일부 모듈은 아직 완전히 오픈되지 않았으며, 커뮤니티 기반 개선이 필요하다.
실용적 활용
Open-AoE는 로봇 개발, 월드 모델 학습, 인간-로봇 전이 연구에 활용 가능하다. 특히, VLA 정책, WAM, World Models 학습에 필요한 구조화된 데이터와 도구체인을 통해 연구자들이 데이터 전처리 및 모델 학습 과정을 효율적으로 수행할 수 있다. 또한, 저비용 스마트폰 캡처를 기반으로 일반 사용자도 데이터 기여에 참여할 수 있어, 대규모 커뮤니티 기반 데이터셋 확장이 가능하다.