ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Xionghao Wu, Yijun Yang, Shiyang Zhou, Haoze Sun, Jianhui Liu, Songsong Yu, Jiyao Zhang, Wenbo Li, Bo Wang, Guoqing Ma, Lin Song, Renjie Liao, Shenghe Zheng, Wei Tang, Xiaojuan Qi, Yanwei Li, Yuan Zhang, Zhuotao Tian, Haoyang Huang, Nan Duan

arXiv:2609.00188 · 2026-09-02 공개 · arXiv · PDF

robot-manipulation generalization world-action-models egocentric-videos real-time-control asynchronous-training zero-shot-evaluation slow-fast-architecture

Abstract

Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of embodied experience, capturing object interactions, contact dynamics, tool use, and long-horizon behaviors across diverse environments. The central challenge is how to convert this abundant but action-free experience into effective robot control. We introduce ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video. ZimaBlue follows a three-stage training curriculum: it first performs causal embodied video pre-training on large-scale human and robot egocentric videos, then grounds the learned visual dynamics in heterogeneous robot trajectories through video-action mid-training with a unified action representation, and finally specializes the model to a target robot for deployment. To make generative WAMs practical for real-time control, ZimaBluefurther adopts an asynchronous Slow-Fast dual-system architecture, where a high-capacity Slow world model provides generalizable spatiotemporal representations and a lightweight Fast branch enables 30 Hz action prediction on NVIDIA RTX 4090. On real-robot zero-shot evaluations, scaling from target-robot data alone to over 120,000 hours of embodied video improves success from 36.1% to 77.8%. ZimaBlue further delivers strong performance across multiple benchmarks, with particularly pronounced gains on unseen tasks.

한국어 요약

한 줄 요약

ZimaBlue는 대규모 영상 데이터를 활용해 일반화 가능한 월드 액션 모델(WAM)을 학습하는 3단계 훈련 프레임워크로, 120,000시간의 영상 학습으로 로봇 성공률을 36.1%에서 77.8%로 향상시킴.

핵심 기여도

핵심 아이디어

ZimaBlue는 로봇 제어를 위한 WAM 학습에서 기존의 액션 라벨이 필요한 로봇 트래젝토리에 의존하지 않고, 대규모 egocentric 영상 데이터를 활용하는 새로운 접근법을 제시한다. 핵심 아이디어는 **비액션 라벨 영상에서 인과적 시각 동역학을 학습한 후, 이동하는 로봇 데이터로 정렬하고, 마지막으로 특정 로봇에 맞춘 후처리를 통해 일반화를 달성**하는 것이다. 이는 VLA 모델이 액션 라벨 데이터에 제한되는 문제를 해결하며, 특히 장기적 행동과 물리적 상호작용을 포함한 복잡한 작업에서 효과적이다. WAM은 **비주얼 변화와 액션의 인과 관계를 학습**하여, 로봇이 시각적 변화를 기반으로 적절한 액션을 예측하도록 유도한다.

기술적 접근법

ZimaBlue는 다음과 같은 3단계 훈련 프로세스를 채택한다:
1. **Causal Embodied Video Pre-training**: 대규모 인간 및 로봇 egocentric 영상(120,000시간 이상)을 사용하여 미래 시각 상태를 예측하는 인과적 학습 수행.
2. **Video-Action Mid-training**: 다양한 로봇 플랫폼에서 수집된 트래젝토리와 영상을 결합, **통일된 액션 표현(Unified Action Representation)** 을 도입하여 이질적인 데이터 정렬.
3. **Target-Robot Post-training**: 특정 로봇(예: 7-DoF Franka arm)에 맞춘 후처리 훈련을 통해 실제 배포에 최적화.

또한, **Slow-Fast 이중 시스템**을 도입:

주요 결과

의의 및 한계

ZimaBlue는 대규모 영상 데이터를 기반으로 로봇 학습을 확장하는 새로운 패러다임을 제시하며, 특히 **장기적 작업과 물리적 상호작용이 필요한 작업에서 실질적인 성능 향상**을 보인다. Slow-Fast 구조는 실시간 제어와 고급 추론을 분리하여, **실용적이고 일반화 가능한 제어 시스템 구축 가능**하다는 점에서 학술적·산업적 가치가 크다.

하지만, **120,000시간의 영상은 전체 잠재력의 일부에 불과**하며, 더 다양한 환경과 로봇 플랫폼에서의 평가가 필요하다. 또한, **비물리적 콘텐츠(예: 애니메이션, 편집된 전환)**가 포함된 영상은 물리적 상호작용 학습에 영향을 줄 수 있으며, 이는 한계로 작용한다.

실용적 활용

ZimaBlue는 **다양한 로봇 플랫폼(예: 7-DoF Franka arm)에서의 장기적 작업 수행**, **실시간 제어가 필요한 산업 자동화**, **다양한 환경에서의 즉석 학습 및 적응**에 활용 가능하다. 특히, **액션 라벨 데이터 수집이 어려운 상황에서 대규모 영상 데이터를 활용한 학습**이 필요한 로봇 개발에 적합하다.