Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou

arXiv:2607.15330 · 2026-07-20 공개 · arXiv · PDF

vision-language-action post-training pre-training data-efficiency mobile-manipulation robodojo real-world-trajectories auto-labeling-pipeline

Abstract

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

한국어 요약

한 줄 요약

Xiaomi-Robotics-1은 10만 시간 이상의 실제 조작 데이터로 학습한 VLA 모델로, RoboCasa365에서 57.6% 성공률을 달성했다.

핵심 기여도

핵심 아이디어

Xiaomi-Robotics-1은 대규모 실제 조작 데이터(10만 시간 이상)를 기반으로 VLA 모델을 학습시켜, 다양한 환경에서 즉시 사용 가능한 제어 능력을 얻는다. 핵심은 **두 단계 학습**(pre-training과 post-training)과 **자동 라벨링 파이프라인**이다. pre-training에서는 UMI 장치를 통해 수집된 데이터로 모델이 일반적인 조작 능력을 습득하게 하며, post-training에서는 실제 로봇의 제어와 인간의 명령어에 맞게 정렬한다. 특히, **Qwen3-VL 기반의 VLM**과 **DiT**(Diffusion Transformer)를 결합한 **Mixture-of-Transformers(MoT)** 아키텍처를 사용하여, 시각-언어 정보를 기반으로 정확한 액션을 생성한다.

기술적 접근법

주요 결과

의의 및 한계

Xiaomi-Robotics-1은 대규모 실제 데이터와 두 단계 학습을 통해 VLA 모델의 성능을 획기적으로 향상시켰으며, 특히 **자동 라벨링 파이프라인**을 통해 높은 품질의 학습 데이터를 확보한 점에서 학술적·실용적 가치가 있다. 또한, **실제 로봇 환경에서 즉시 사용 가능한 성능**을 보여주며, 기존 모델 대비 높은 데이터 효율성을 나타냈다. 그러나 한계로는 **UMI 장치에 의존하는 데이터 수집 방식**이 일반화에 한계가 있을 수 있으며, **실제 로봇 환경에서의 장기적 안정성**에 대한 평가가 부족하다는 점이 언급된다.

실용적 활용

Xiaomi-Robotics-1은 서비스 로봇, 산업 자동화, 가정용 로봇 등 다양한 분야에서 즉시 사용 가능한 제어 정책을 제공할 수 있다. 특히, **최소한의 데이터로도 새로운 작업에 빠르게 적응**할 수 있어, 로봇 개발 및 배포 과정에서 데이터 수집 비용을 줄이는 데 유용하다.