RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Min Xu, Lunkai Lin, Zhiqiang Xie, Mingyu Ding, Ping Luo

arXiv:2504.13059 · 2026-07-27 공개 · arXiv · PDF

large-language-models code-generation sim-to-real object-manipulation cobot-magic spatial-relation-aware dual-arm-robotics generative-digital-twins

Abstract

In the rapidly advancing field of robotics, dual-arm co-ordination and complex object manipulation are essential capabilities for developing advanced autonomous systems. However, the scarcity of diverse, high-quality demonstration data and real-world-aligned evaluation benchmarks severely limits such development. To address this, we introduce RoboTwin, a generative digital twin framework that uses 3D generative foundation models and large language models to produce diverse expert datasets and provide a real-world-aligned evaluation platform for dual-arm robotic tasks. Specifically, RoboTwin creates varied digital twins of objects from single 2D images, generating realistic and interactive scenarios. It also introduces a spatial relation-aware code generation framework that combines object annotations with large language models to break down tasks, determine spatial constraints, and generate precise robotic movement code. Our framework offers a comprehensive benchmark with both simulated and real-world data, enabling standardized evaluation and better alignment between simulated training and real-world performance. We validated our approach using the open-source COBOT Magic Robot platform. Policies pre-trained on RoboTwin-generated data and fine-tuned with limited real-world samples demonstrate significant potential for enhancing dual-arm robotic manipulation systems by improving success rates by over 70% for single-arm tasks and over 40% for dual-arm tasks compared to models trained solely on real-world data.

한국어 요약

한 줄 요약

RoboTwin은 3D 생성 모델과 대형 언어 모델을 결합한 디지털 트윈 프레임워크로, 이중 팔 로봇 조작 성능을 70% 이상 향상시킨다.

핵심 기여도

핵심 아이디어

기존 이중 팔 로봇 학습은 고비용·시간 소요의 인간 텔레오퍼레이션에 의존하거나, 시뮬레이션에서 생성된 고정된 데이터에 제한되어 일반화 능력이 낮았다. RoboTwin은 이 문제를 해결하기 위해 3D 생성 기초 모델과 대형 언어 모델을 결합한 디지털 트윈 기반 학습 프레임워크를 제안한다. 2D RGB 이미지에서 3D 객체를 생성하고, 객체의 기능적 축, 접근 축, 접촉점 등 공간 어노테이션을 생성하여, 대형 언어 모델이 이를 바탕으로 정확한 로봇 동작 코드를 생성한다. 이는 단순한 시뮬레이션 데이터 생성을 넘어, 실제 세계와 일치하는 정확한 공간 제약을 반영한 실행 가능한 코드 생성을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

RoboTwin은 이중 팔 로봇 학습의 데이터 수집 및 평가 문제를 해결하며, 시뮬레이션-실제 간의 간극을 효과적으로 줄인다. 특히, 3D 생성 모델과 대형 언어 모델의 결합은 실제 세계와 일치하는 정확한 작업 코드 생성을 가능하게 하며, 이는 기존 시뮬레이션 기반 접근법보다 훨씬 유연하고 확장 가능하다. 그러나 이중 팔 조작의 복잡성은 여전히 한계를 드러내며, 더 복잡한 작업에 대한 알고리즘 개선이 필요하다.

실용적 활용

RoboTwin은 제조, 의료, 물류 등에서 이중 팔 로봇의 정밀한 작업 수행을 가능하게 하며, 실제 세계와 시뮬레이션 간의 정책 전이를 용이하게 한다. 특히, 고비용 텔레오퍼레이션 대신 저비용으로 훈련 데이터를 생성할 수 있어 산업 현장에서 즉각적인 활용이 가능하다.