Masked Visual Actions for Unified World Modeling

Hadi Alzayer, Wenlong Huang, Haonan Chen, Christopher Luey, Lvmin Zhang, Maneesh Agrawala, Gordon Wetzstein, Li Fei-Fei, Yilun Du, Jiajun Wu, Jia-Bin Huang

arXiv:2607.19343 · 2026-07-22 공개 · arXiv · PDF

policy-evaluation visual-fidelity forward-dynamics model-based-planning inverse-modeling real-video-finetuning masked-visual-actions robotic-world-modeling

Abstract

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.

한국어 요약

한 줄 요약

Masked Visual Actions는 비디오 모델에 시각적 액션을 주입하여 로봇 월드 모델링을 통합하는 새로운 제어 인터페이스이다.

핵심 기여도

핵심 아이디어

기존 로봇 월드 모델은 텍스트, 트랙, 포스, 키포인트 등 비시각적 신호로 조건을 주입하지만, 이는 모델의 사전 학습된 시각 경험과 미스얼라인되어 있다. Masked Visual Actions는 액션을 비디오 내 임의 엔티티의 부분적으로 드러난 궤적으로 표현함으로써, 모델의 픽셀 공간 내에서 직접적으로 액션을 표현한다. 이 방식은 로봇 움직임을 드러내면 정방향 역학 모델로, 객체 움직임을 드러내면 역모델로 동작하게 한다. 이는 동일한 모델을 사용하여 정방향 및 역방향 추론을 수행할 수 있음을 의미하며, 이는 기존의 별도 아키텍처에 의존하는 접근과 구별된다.

기술적 접근법

주요 결과

의의 및 한계

Masked Visual Actions는 비디오 모델의 사전 학습된 경험을 효과적으로 활용하여, 로봇 월드 모델링의 두 방향(정방향 및 역방향)을 동일한 모델 내에서 처리할 수 있는 새로운 인터페이스를 제시한다. 이는 로봇 정책 평가, 계획, 역모델링 등 다양한 응용에서 실용성을 증가시키며, 학습 비용을 낮출 수 있다. 그러나 모델은 객체 상호작용의 상관관계를 학습할 뿐 인과관계는 아님을 명시하며, 기반 비디오 모델의 표현력과 추론 속도에 제한을 받는다.

실용적 활용

Masked Visual Actions는 로봇 정책 평가, 기반 모델 기반의 계획, 역모델링 등 다양한 로봇 조작 상황에서 활용 가능하다. 특히, 실제 환경에서의 학습 데이터가 제한적일 때, 시뮬레이션과 실제 데이터를 결합한 마스킹 예제로 학습 가능한 점에서 유용하며, 로봇 개발 비용을 절감하고 접근성을 높일 수 있다.