Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li

arXiv:2607.14183 · 2026-07-21 공개 · arXiv · PDF

world-models robot-learning egocentric-vision hand-reconstruction manipulation-dataset embodied-learning camera-trajectory vla-policies

Abstract

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.

한국어 요약

한 줄 요약

Open-AoE는 스마트폰으로 수집된 2,000시간의 제1인칭 조작 영상과 처리 파이프라인, 학습 도구를 제공하는 오픈 인프라이다.

핵심 기여도

핵심 아이디어

기존 제1인칭 데이터셋은 스케일 확장, 구조화된 어노테이션, 재사용 가능한 도구체인을 결합하지 못했다. Open-AoE는 스마트폰 기반 저비용 캡처와 처리, 재사용 가능한 학습 인프라를 통합한 체계적인 솔루션을 제시한다. 핵심 아이디어는 "capture-process-reconstruct-train"이라는 데이터 생산 루프를 오픈 소스로 제공하는 것이다. 이는 VLA 정책, WAM, World Models 등 다양한 모델 학습에 직접적으로 활용할 수 있도록 데이터를 구조화하고, MANO 기반 손 포즈, 카메라 궤적, 시간 분할된 원자적 행동 등 세부 정보를 포함한다.

기술적 접근법

주요 결과

의의 및 한계

Open-AoE는 제1인칭 데이터셋의 확장성, 구조화, 재사용성을 동시에 해결한 첫 오픈 인프라로, 로봇 학습과 월드 모델링에 실질적 기여를 할 수 있다. 특히, 스마트폰 기반 저비용 캡처와 오픈 도구체인은 데이터 기여와 활용의 장벽을 낮춘다. 그러나 현재 데이터셋은 특정 환경과 장치에 의존적이며, 더 넓은 범위의 장치와 상황에서의 데이터 수집이 필요하다. 또한, 도구체인의 일부 모듈은 아직 완전히 오픈되지 않았으며, 커뮤니티 기반 개선이 필요하다.

실용적 활용

Open-AoE는 로봇 개발, 월드 모델 학습, 인간-로봇 전이 연구에 활용 가능하다. 특히, VLA 정책, WAM, World Models 학습에 필요한 구조화된 데이터와 도구체인을 통해 연구자들이 데이터 전처리 및 모델 학습 과정을 효율적으로 수행할 수 있다. 또한, 저비용 스마트폰 캡처를 기반으로 일반 사용자도 데이터 기여에 참여할 수 있어, 대규모 커뮤니티 기반 데이터셋 확장이 가능하다.