GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
GigaBrain Team, Angen Ye, Axiang Sun, Can Jin, Chenxi Cheng, Chong Shi, Dengke Shang, Dingqian Zhang, Guan Huang, Guangqiang Wang, Guangqing Ding, Guo Li, Hangcong Li, Hengyu Zhong, Hongtao Lu, Jianbo Qin, Jiming Mao, Jing Zhu, Jindi Lv, Jingzhi Cui, Junjie Xie, Junyi Bao, Kai Liu, Lei Yuan, Limin Long, Lv Feng, Mingming Yu, Peng Li, Pengfei Yi, Qi Li, Qianli Zhang, Qingfang Li, Qitang Hu, Rui Zhang, Shaoyan Sun, Shibo Sun, Shiying Duan, Tenghui Chen, Tianze Liu, Weijie Ke, Wenyao Xue, Xiaofeng Wang, Xiaoyu Tian, Xinyu Liu, Xinze Chen, Yang Wang, Yankai Wang, Yejun Zeng, Yifan Li, Yifei Nie, Yilong Li, Yilong Liu, Yongchao Feng, Yumeng Wang, Yun Ye, Zhichao Liu, Ziheng He, Zonghai Yang, Zheng Zhu
arXiv:2608.15875 · 2026-08-26 공개 · arXiv · PDF
foundation-models vision-language-action instruction-following embodied-agents multi-embodiment embodied-foundation-models robot-embodiments alignment-training
Abstract
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including π_{0.5}, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
한국어 요약
한 줄 요약
GigaBrain-0.7는 37,000시간 이상의 이질적 로봇 데이터로 사전 학습한 3시스템 아키텍처 기반의 VLA 모델로, 다양한 로봇 형태에서의 제로샷 성능과 작업 성공률을 대폭 향상시킨다.
핵심 기여도
- 37,000시간 이상의 이질적 로봇 데이터로 사전 학습하여 이전 모델 대비 더 넓은 범용성 확보
- System 1, 2, 3의 3시스템 아키텍처로 이해, 예측, 실행을 통합
- One-stage alignment training으로 시각-언어 이해와 다형태 행동 생성을 동시에 최적화
- GigaBrain-0 시리즈 대비 제로샷, 언어 조건 실행, 후처리 작업 성공률에서 10~20% 개선
핵심 아이디어
GigaBrain-0.7은 단순히 더 많은 데이터를 학습하는 것을 넘어, 이질적인 로봇 데이터를 효과적으로 통합하고, 이해, 예측, 실행을 분리하지 않고 통합하는 3시스템 아키텍처를 도입했다. System 1은 로봇의 제어 신호를 생성하는 Action Expert(0.5B)를 포함하며, System 2는 PaliGemma2(3B)를 기반으로 시각-언어 이해와 계획을 담당하고, System 3은 GigaWorld-1(5B) 기반의 월드 모델을 통해 미래 상태를 예측하고 가치 평가를 수행한다. 이는 기존의 반응형(observation-to-action) 접근에서 벗어나, 장기적 계획과 예측을 통한 행동 개선을 가능하게 한다.
기술적 접근법
- **모델 아키텍처**: System 1(Action and Control), System 2(Understanding and Planning), System 3(Prediction and Evaluation)의 3시스템 구성
- **데이터**: 37,000시간 이상의 이질적 로봇 데이터, 16개 로봇 형태, 약 270백만 개의 시각-언어 샘플
- **사전 학습**: LeRobot v3.0 포맷 변환, LLM 기반 언어 지시 수정, 다단계 품질 검증을 통한 데이터 정제
- **학습 전략**: One-stage alignment training, Soft Knowledge Insulation(Soft KI)을 사용하여 VLM 핵심 능력을 유지하면서 로봇 제어에 적응
- **후처리 학습**: System 3은 고정되고, System 1과 2는 공동 최적화
주요 결과
- **Maker H01 및 AgileX PiPER/PiPER-X 플랫폼**에서 제로샷, 다태스크 실행, 장기 조작 성능 향상
- **System 3의 SubImage+Value 조건**에서 의류 접기 작업 성공률 100% 유지, 평균 점수 88.3% 달성 (기본 모델 대비 +20%)
- **선물 포장 작업**에서 SubImage+Value 조건 시 성공률 80% (기본 모델 0%)
- **큐브 정렬 작업**에서 SubImage+Value 조건 시 성공률 55% (기본 모델 40%)
- **GigaBrain-0 시리즈 대비** 제로샷, 언어 조건 실행, 후처리 작업 성공률에서 10~20% 개선
의의 및 한계
GigaBrain-0.7는 이질적인 로봇 데이터를 통합하고, 장기적 계획과 예측을 통한 행동 개선을 가능하게 함으로써, 기존 VLA 모델의 반응형 제한을 극복했다는 점에서 학술적 의의가 있다. 특히, System 3의 미래 상태 예측과 가치 평가가 작업 진행의 효과성을 높이는 데 기여한 것으로 나타났다. 그러나 모든 작업에서 100% 성공률을 달성하지 못했으며, 일부 작업에서는 OOD(Out-of-Distribution) 환경에서의 일반화 능력이 제한된 것으로 보인다. 또한, 37,000시간 이상의 데이터를 처리하는 데 필요한 계산 자원과 저장 공간은 여전히 큰 과제이다.
실용적 활용
GigaBrain-0.7는 다양한 로봇 형태(이중 팔 조작기, 인간형 로봇 등)에서의 제로샷 실행과 장기적 작업 수행에 적합하며, 가정용 및 산업용 로봇 제어, 언어 기반 작업 지시 실행 등에 활용될 수 있다. 특히, AgileX PiPER와 Maker H01 플랫폼에서의 성능 개선은 실제 산업 현장에서의 적용 가능성을 높인다.