GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

GigaBrain Team, Angen Ye, Axiang Sun, Can Jin, Chenxi Cheng, Chong Shi, Dengke Shang, Dingqian Zhang, Guan Huang, Guangqiang Wang, Guangqing Ding, Guo Li, Hangcong Li, Hengyu Zhong, Hongtao Lu, Jianbo Qin, Jiming Mao, Jing Zhu, Jindi Lv, Jingzhi Cui, Junjie Xie, Junyi Bao, Kai Liu, Lei Yuan, Limin Long, Lv Feng, Mingming Yu, Peng Li, Pengfei Yi, Qi Li, Qianli Zhang, Qingfang Li, Qitang Hu, Rui Zhang, Shaoyan Sun, Shibo Sun, Shiying Duan, Tenghui Chen, Tianze Liu, Weijie Ke, Wenyao Xue, Xiaofeng Wang, Xiaoyu Tian, Xinyu Liu, Xinze Chen, Yang Wang, Yankai Wang, Yejun Zeng, Yifan Li, Yifei Nie, Yilong Li, Yilong Liu, Yongchao Feng, Yumeng Wang, Yun Ye, Zhichao Liu, Ziheng He, Zonghai Yang, Zheng Zhu

arXiv:2608.15875 · 2026-08-26 공개 · arXiv · PDF

foundation-models vision-language-action instruction-following embodied-agents multi-embodiment embodied-foundation-models robot-embodiments alignment-training

Abstract

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.

한국어 요약

한 줄 요약

GigaBrain-0.7는 37,000시간 이상의 이질적 로봇 데이터로 사전 학습한 3시스템 아키텍처 기반의 VLA 모델로, 다양한 로봇 형태에서의 제로샷 성능과 작업 성공률을 대폭 향상시킨다.

핵심 기여도

핵심 아이디어

GigaBrain-0.7은 단순히 더 많은 데이터를 학습하는 것을 넘어, 이질적인 로봇 데이터를 효과적으로 통합하고, 이해, 예측, 실행을 분리하지 않고 통합하는 3시스템 아키텍처를 도입했다. System 1은 로봇의 제어 신호를 생성하는 Action Expert(0.5B)를 포함하며, System 2는 PaliGemma2(3B)를 기반으로 시각-언어 이해와 계획을 담당하고, System 3은 GigaWorld-1(5B) 기반의 월드 모델을 통해 미래 상태를 예측하고 가치 평가를 수행한다. 이는 기존의 반응형(observation-to-action) 접근에서 벗어나, 장기적 계획과 예측을 통한 행동 개선을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

GigaBrain-0.7는 이질적인 로봇 데이터를 통합하고, 장기적 계획과 예측을 통한 행동 개선을 가능하게 함으로써, 기존 VLA 모델의 반응형 제한을 극복했다는 점에서 학술적 의의가 있다. 특히, System 3의 미래 상태 예측과 가치 평가가 작업 진행의 효과성을 높이는 데 기여한 것으로 나타났다. 그러나 모든 작업에서 100% 성공률을 달성하지 못했으며, 일부 작업에서는 OOD(Out-of-Distribution) 환경에서의 일반화 능력이 제한된 것으로 보인다. 또한, 37,000시간 이상의 데이터를 처리하는 데 필요한 계산 자원과 저장 공간은 여전히 큰 과제이다.

실용적 활용

GigaBrain-0.7는 다양한 로봇 형태(이중 팔 조작기, 인간형 로봇 등)에서의 제로샷 실행과 장기적 작업 수행에 적합하며, 가정용 및 산업용 로봇 제어, 언어 기반 작업 지시 실행 등에 활용될 수 있다. 특히, AgileX PiPER와 Maker H01 플랫폼에서의 성능 개선은 실제 산업 현장에서의 적용 가능성을 높인다.