InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

arXiv:2609.31394 · 2026-09-28 공개 · arXiv · PDF

robot-manipulation world-action-models mixture-of-transformers visual-dynamics large-scale-pretraining umi-data ego2robot scene-semantics

Abstract

World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We introduce InternW0-Δ, a unified WAM pretrained on a heterogeneous corpus that outperforms prior methods across simulation benchmarks and real-robot platforms. InternW0-Δ combines pretrained visual dynamics, scene-level semantics, 4D geometric and motion priors, and action generation within a Mixture-of-Transformers (MoT) framework. A pretrained video expert and an action expert interact under semantic guidance from a frozen VLM, while a pretrained 4D foundation model injects geometric and motion priors through training-only distillation. We further introduce Causal Imprint, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference. For large-scale joint training, we construct a heterogeneous corpus of robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data, curated and aligned under a common state-action representation. The resulting corpus contains over 20K hours of processed training data, to our knowledge the largest open-source corpus of its kind. We pretrain InternW0-Δ on this corpus and demonstrate strong performance across simulation benchmarks and real-robot platforms. We will open source the training code, model weights, infrastructure, data-processing pipeline, and processed data where licenses permit. Project page: https://internrobotics.github.io/InternW0-Delta/

한국어 요약

한 줄 요약

InternW0-Δ는 20K시간 이상의 이질적 데이터로 학습한 세계 행동 모델로, 시각 역학과 행동 생성을 통합하여 로봇 제어 성능을 향상시킨다.

핵심 기여도

핵심 아이디어

InternW0-Δ는 대규모 사전 학습 모델의 다양한 사전 지식을 통합하여 로봇 행동 생성을 최적화하는 새로운 접근법을 제시한다. 기존 방법에서는 시각 역동성 예측과 행동 생성을 별도로 처리하거나, 미래 영상 생성을 필요로 하는 경우가 많았지만, InternW0-Δ는 Causal Imprint를 통해 미래 영상 없이도 예측적 표현을 학습한다. 이는 추론 시 미래 영상 롤아웃 없이 행동을 직접 생성할 수 있게 한다. 또한, Track4World 모델을 활용한 4D-aware distillation을 통해 기하학적 및 운동학적 사전 지식을 주입함으로써, 행동 전문가가 더 정확한 제어를 수행할 수 있도록 지원한다. MoT 아키텍처는 이러한 전문가 간의 정보 흐름을 효율적으로 관리하며, VLM을 통해 작업 조건에 맞는 의미 정보를 제공한다.

기술적 접근법

주요 결과

의의 및 한계

InternW0-Δ는 대규모 이질적 데이터를 기반으로 한 통합된 행동 생성 모델로서, 다양한 로봇 플랫폼과 작업 환경에서 뛰어난 성능을 보인다. 특히, 미래 영상 없이도 행동을 생성할 수 있는 Causal Imprint는 실시간 제어에 유리하며, 4D-aware distillation은 기하학적 이해를 강화하여 복잡한 작업에서도 안정적인 성능을 제공한다. 그러나 모델의 복잡성은 훈련 및 배포 시 높은 계산 자원을 요구하며, 특정 작업에 대한 세부 조정이 필요할 수 있다. 또한, 데이터셋의 이질성은 모델의 일반화 능력을 향상시키지만, 일부 작업에서는 데이터 부족으로 인한 성능 저하가 발생할 수 있다.

실용적 활용

InternW0-Δ는 다양한 로봇 플랫폼 (예: 그리퍼, 다exterous-hand)에서의 일반적 제어, 복잡한 작업 수행 (예: 이동형 이중 조작), 분산 환경에서의 실시간 제어 등에 적용 가능하다. 특히, 미래 영상 생성 없이 행동을 예측할 수 있는 구조는 실시간 성능 향상에 유리하며, 대규모 데이터셋과 공개된 인프라를 통해 연구 및 산업 현장에서의 활용이 확대될 수 있다.