RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

Yuheng Ji, Huajie Tan, Jiayu Shi, Xiaoshuai Hao, Yuan Zhang, Hengyuan Zhang, Pengwei Wang, Mengdi Zhao, Yao Mu, Pengju An, Xinda Xue, Qi Su, Huaihai Lyu, Xiaolong Zheng, Jiaming Liu, Zhongyuan Wang, Shanghang Zhang

arXiv:2502.21257 · 2026-07-27 공개 · arXiv · PDF

multimodal-llm robotic-manipulation long-video trajectory-prediction high-resolution-images task-planning affordance-perception sharerobot-dataset

Abstract

Recent advancements in Multimodal Large Language Models (MLLMs) have shown remarkable capabilities across various multimodal contexts. However, their application in robotic scenarios, particularly for long-horizon manipulation tasks, reveals significant limitations. These limitations arise from the current MLLMs lacking three essential robotic brain capabilities: Planning Capability, which involves decomposing complex manipulation instructions into manageable sub-tasks; Affordance Perception, the ability to recognize and interpret the affordances of interactive objects; and Trajectory Prediction, the foresight to anticipate the complete manipulation trajectory necessary for successful execution. To enhance the robotic brain’s core capabilities from abstract to concrete, we introduce ShareRobot, a high-quality heterogeneous dataset that labels multi-dimensional information such as task planning, object affordance, and end-effector trajectory. ShareRobot’s diversity and accuracy have been meticulously refined by three human annotators. Building on this dataset, we developed RoboBrain, an MLLM-based model that combines robotic and general multi-modal data, utilizes a multi-stage training strategy, and incorporates long videos and high-resolution images to improve its robotic manipulation capabilities. Extensive experiments demonstrate that RoboBrain achieves state-of-the-art performance across various robotic tasks, highlighting its potential to advance robotic brain capabilities. Project website: RoboBrain.

한국어 요약

한 줄 요약

RoboBrain은 ShareRobot 데이터셋을 기반으로, 로봇 조작의 추상적 지시를 구체적 행동으로 변환하는 통합 MLLM 모델로, 기존 모델 대비 14.6~18.75% 성능 향상.

핵심 기여도

핵심 아이디어

기존 MLLM은 로봇 조작에서 필요한 세 가지 핵심 능력—플래닝, affordance 인식, 트레젝토리 예측—을 갖추지 못한다. 예를 들어, 물병을 들어 올리고 물을 따르는 작업은 "접근 및 들어 올리기", "물병 입구를 컵 위로 이동", "기울여 따르기"와 같은 하위 작업으로 분해되어야 하며, 각 작업에 대해 객체의 잡을 수 있는 영역(affordance)을 인식하고, 엔드이펙터의 정확한 이동 경로를 예측해야 성공한다. RoboBrain은 이러한 능력을 강화하기 위해 ShareRobot 데이터셋을 기반으로 설계되었으며, LLaVA 기반의 기초 모델과 A-LoRA, T-LoRA 모듈을 결합하여 각각 affordance 인식과 트레젝토리 예측을 담당한다.

기술적 접근법

주요 결과

의의 및 한계

RoboBrain은 추상적 지시를 구체적 로봇 행동으로 변환하는 데 있어 기존 MLLM의 한계를 극복한 첫 사례로, 로봇 학습 및 자율 조작 분야에 중요한 발전을 의미한다. 특히, ShareRobot 데이터셋은 로봇 학습의 정확성과 다양성을 동시에 확보한 고질량 데이터셋으로, 향후 연구에 기반 자료로 활용될 수 있다. 그러나 현재 모델은 특정 환경에서만 훈련되었으며, 실시간 성능이나 다중 로봇 협업 환경에서의 확장성은 추가 연구가 필요하다.

실용적 활용

RoboBrain은 산업 로봇, 서비스 로봇, 의료 로봇 등에서 복잡한 조작 작업을 자동화하는 데 활용 가능하다. 예를 들어, 물류 시스템에서 상자 정렬, 의료 환경에서 의료 장비 조작, 가정용 로봇에서 물건 이동 등 다양한 분야에서 적용 가능하다.