Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Ye Wang, Pei Lin, Xiong-Hui Chen, Haoqi Yuan, Zhixuan Liang, Yiyang Huang, Anzhe Chen, Zixing Lei, Jie Zhang, Tao Zhang, Haoyang Li, Tong Zhang, Chenxi Xiao, Ziyuan Jiao, Qin Jin

arXiv:2608.02580 · 2026-08-06 공개 · arXiv · PDF

vision-language-action dataset-curation egocentric-videos generalization-evaluation robot-morphology robot-pretraining ego2robot robot-data-synthesis

Abstract

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-language-action models at scale remains unexplored. We present Ego2Robot, a scalable pipeline that converts egocentric human manipulation videos into robot training data through action retargeting, robot-arm visual synthesis, and multi-level quality curation. Ego2Robot supports both curated datasets and in-the-wild videos, producing 18,561 hours of robot training data spanning 15 robot morphologies, making it the largest ego-to-robot dataset to date. To evaluate generalization, we extend RoboTwin2.0 with disentangled perturbation axes covering visual appearance, scene layout, embodiment morphology, and task semantics. Experiments show that joint pretraining on Ego2Robot-synthesized and robot data consistently improves out-of-distribution generalization across multiple perturbation types, with benefits validated on real-robot deployment. Project page: https://www-ye.github.io/ego2robot_blog/

한국어 요약

한 줄 요약

Ego2Robot은 인간 제1인칭 조작 영상에서 로봇 학습 데이터를 대규모로 생성하는 파이프라인으로, 18,561시간의 데이터를 생성해 일반화 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 로봇 학습 데이터 수집은 비용이 높고 다양성이 제한적이지만, 인간 제1인칭 조작 영상은 대규모로 수집 가능하며 풍부한 작업 정보를 포함한다. 그러나 인간과 로봇의 체형 차이로 인해 직접 변환은 어려움. Ego2Robot은 액션 리타겟팅과 시각 정렬을 통해 인간 영상에서 로봇 학습 데이터를 생성하며, 품질 큐레이션을 통해 데이터 신뢰도를 높인다. 이는 로봇 정책의 일반화 능력을 향상시키는 새로운 패러다임을 제시한다. 핵심 통찰은 인간 조작 영상에 내재된 상호작용 규칙이 로봇 학습에 유용하다는 점이다.

기술적 접근법

주요 결과

의의 및 한계

Ego2Robot은 대규모 로봇 학습 데이터 생성을 위한 새로운 접근법으로, 인간 영상에서 로봇 정책의 일반화 능력을 향상시키는 데 기여한다. 특히, 시각적 불확실성, 체형 차이, 의미적 변형에 대한 내성을 높인다. 그러나 인간과 로봇의 체형 차이로 인한 도메인 간 격차는 여전히 존재하며, 이는 일부 퍼트urbation에서 성능 향상이 제한적일 수 있음을 시사한다. 또한, 합성 데이터의 품질은 품질 큐레이션 단계에 크게 의존한다.

실용적 활용

Ego2Robot은 로봇 제조사, 연구소, 교육 기관 등에서 대규모 학습 데이터 생성에 활용 가능하다. 특히, 실제 로봇 데이터 수집이 어려운 환경에서 멀티태스크 학습 및 일반화 능력 향상에 유용하며, 저비용으로 로봇 정책 개발을 가속화할 수 있다.