OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

Yuran Wang, Siqiao Huang, Mingleyang Li, Chenhao Zhang, Jiaqi Liang, Weiyang Jin, Yue Chen, Xuemin Chi, Donghao Zhou, Qize Yu, Yu-Kai Wang, Yuhan Rui, Shenzhe Yao, Zhen Yuan, Zhenhao Shen, Kefei Zhu, Zijie Zhu, Ning Gao, Xiaowei Chi, Guanqi He, Shanghang Zhang, Hao Dong, Lin Shao, Hang Zhao

arXiv:2609.07398 · 2026-09-09 공개 · arXiv · PDF

robot-manipulation latent-space world-action-models real-world-evaluation simulation-benchmarks egocentric-data embodied-pretraining modular-pretraining

Abstract

World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-α, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-α delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.

한국어 요약

한 줄 요약

OpenWAM은 모듈화된 연구 스택을 통해 시스템적인 월드-액션 모델 사전학습을 구현한 오픈 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 월드-액션 모델은 생성 백본, 시각 표현, 아키텍처 등이 강하게 결합되어 있어 설계 요소 간의 상호작용을 분석하기 어려웠다. OpenWAM은 이를 모듈화하여 실험적 제어가 가능하도록 구조화함으로써, 어떤 설계 선택이 중요한지를 명확히 파악할 수 있도록 했다. 핵심 통찰 중 하나는 "체화 사전학습이 도메인 외 일반화를 향상시킨다"는 점이며, 이는 사전학습 데이터가 체화적(egocentric)인 경우에 특히 효과적임을 보여준다.

기술적 접근법

주요 결과

의의 및 한계

OpenWAM은 월드-액션 모델 연구를 체계적이고 모듈화된 방향으로 전환시킬 수 있는 기반을 제공하며, 다양한 체화 데이터를 통합하는 데 기여한다. 그러나 사용된 데이터셋의 구체적인 구성이나 하이퍼파라미터는 명시되지 않아 재현성 측면에서 한계가 있을 수 있다. 또한, 특정 도메인에서의 성능 개선 폭이 명확히 제시되지 않았다.

실용적 활용

OpenWAM은 로봇 제어, 자율 시스템 개발, 인간-로봇 상호작용 연구 등 체화 인지와 실행이 필요한 분야에서 활용 가능하다. 특히, 다양한 로봇 플랫폼에서의 일반화 능력을 향상시키는 데 유용할 것으로 기대된다.