MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

Ting Huang, Yue Huang, Zeyu Zhang, Shuicheng Yan, Hao Tang

arXiv:2609.06251 · 2026-09-15 공개 · arXiv · PDF

chain-of-thought vision-language-action action-decoder unitree-go2 quadruped-control humanoid-manipulation vln-ce mobile-robotics

Abstract

Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semantic reasoning and low-level locomotion and manipulation control. Existing approaches often rely on implicit reasoning or monolithic action prediction, making it difficult to maintain coherent long-horizon decision making while producing precise and adaptable robot actions. To address this challenge, we propose MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control. The framework learns multi-granularity reasoning over embodied trajectories through supervised Chain-of-Thought (CoT) alignment and reinforcement learning, improving reasoning-to-action consistency beyond purely behavioral supervision. To support both locomotion and manipulation, we further introduce a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers. This design provides a unified perception-reasoning-action interface while decoupling high-level action generation from robot-specific actuation. We conduct extensive evaluations on language-guided navigation, quadruped control, and humanoid mobile manipulation, covering VLN-CE, QUARD, and real-world deployments on Unitree Go2 and G1 robots. MobileVLA-R1 2.0 consistently outperforms strong VLA baselines, achieving an average 1.6 point improvement in SR on VLN-CE and a 10.0 point improvement in full-task success on real-world G1 mobile manipulation tasks over MobileVLA-R1, while demonstrating robust long-horizon instruction following and closed-loop execution across different robotic platforms.

한국어 요약

한 줄 요약

MobileVLA-R1 2.0은 강화학습을 결합한 VLA 프레임워크로, 언어-비전-작업 실행 일관성을 향상시킨다.

핵심 기여도

핵심 아이디어

기존 VLA 시스템은 고수준 추론과 저수준 제어 간의 일관성을 유지하기 어려웠다. MobileVLA-R1 2.0은 CoT 정렬과 강화학습을 결합하여, 로봇의 실제 이동 궤적에 기반한 다중 그레인리티 추론을 학습한다. 이는 단순 행동 감독을 넘어, 추론과 액션 사이의 일관성을 향상시킨다. 또한, 추론 조건화된 액션 디코더를 통해 다중 모달 추론 표현을 작업 수준의 액션 목표로 매핑한 후, 로봇 제어기에서 구체적인 명령으로 변환한다. 이는 고수준 액션 생성과 로봇 특화 제어를 분리하여, 다양한 로봇 플랫폼에 유연하게 적용할 수 있게 한다.

기술적 접근법

주요 결과

의의 및 한계

MobileVLA-R1 2.0은 VLA 시스템에서 추론과 액션 일관성을 향상시키는 새로운 프레임워크로, 실제 로봇 제어에 유용한 인터페이스를 제공한다. 특히, CoT 정렬과 강화학습의 결합은 기존 행동 감독 방식을 넘어선 새로운 학습 전략을 제시한다. 그러나, 실제 세계 환경에서의 일반화 능력이나, 다양한 언어 지시에 대한 처리 능력은 추가 실험을 통해 검증이 필요하다.

실용적 활용

이 연구는 언어 지시를 바탕으로 이동 및 조작을 수행하는 모바일 로봇에 적용 가능하다. 특히, 로봇 제어 시스템 설계, 서비스 로봇, 산업 자동화 분야에서 유용하게 활용될 수 있다.