ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Ran Cheng, Yaxin Peng, Chaomin Shen, Feifei Feng

arXiv:2502.14420 · 2026-07-27 공개 · arXiv · PDF

vision-language-action robot-manipulation mixture-of-experts parameter-efficient vla robot-control multimodal-understanding mmstar

Abstract

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Through a systematic analysis of existing training paradigms in vision-language-action models (VLA), we identify two key challenges: spurious forgetting, where robot training overwrites crucial visual-text alignments, and task interference, where competing control and understanding tasks degrade performance when trained jointly. To overcome these limitations, we propose ChatVLA, a novel framework featuring Phased Alignment Training, which incrementally integrates multimodal data after initial control mastery, and a Mixture-of-Experts architecture to minimize task interference. ChatVLA demonstrates competitive performance on visual question-answering datasets and significantly surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks. Notably, it achieves a six times higher performance on MMMU and scores 47.2% on MMStar with a more parameter-efficient design than ECoT. Furthermore, ChatVLA demonstrates superior performance on 25 real-world robot manipulation tasks compared to existing VLA methods like OpenVLA. Our findings highlight the potential of our unified framework for achieving both robust multimodal understanding and effective robot control.

한국어 요약

한 줄 요약

ChatVLA는 Phased Alignment Training과 Mixture-of-Experts를 결합해 MMMU에서 6배, MMStar에서 47.2% 성능을 달성한 통합형 VLA 모델이다.

핵심 기여도

핵심 아이디어

기존 VLA 모델은 robot control 학습이 visual-text alignment를 덮어써 multimodal 이해력을 약화시키는 "spurious forgetting" 문제와, control과 understanding이 공유 파라미터 공간에서 간섭하는 "task interference" 문제가 발생한다. ChatVLA는 이 두 문제를 해결하기 위해 Phased Alignment Training이라는 2단계 학습 전략을 제안한다. 첫 단계에서 로봇 제어를 먼저 학습한 후, 두 번째 단계에서 multimodal 데이터를 점진적으로 통합하여 alignment를 복구한다. 또한, MoE 아키텍처를 도입해 attention layer는 공유하면서 MLP layer는 분리하여 task 간 간섭을 줄인다. 이는 인지 이론(Dual Coding Theory)에 기반하며, 로봇 제어와 언어-시각 이해가 서로 다른 시스템에서 처리되어야 한다는 통찰을 반영한다.

기술적 접근법

주요 결과

의의 및 한계

ChatVLA는 multimodal understanding과 robot control을 하나의 모델에서 통합하는 데 성공해, 기존 VLA의 단점을 극복한 사례로 주목받는다. 특히, Phased Alignment Training과 MoE 아키텍처는 task 간 간섭과 alignment 손실 문제를 해결하는 데 기여하며, 실제 로봇 조작에서의 성능도 입증되었다. 그러나, VLM backbone의 파라미터 수가 줄어들었음에도 성능이 유지되는 메커니즘에 대한 심층 분석이 부족하며, 더 다양한 환경에서의 일반화 능력 검증이 필요하다. 또한, 실제 로봇 데이터와 visual-text 데이터의 비율 조정에 따른 성능 변화에 대한 정량적 분석도 추가 연구 주제로 제시된다.

실용적 활용

ChatVLA는 실제 가정 환경(욕실, 주방, 테이블탑)에서의 로봇 조작, 물체 인식, 대화형 인터페이스 등에 적용 가능하다. 특히, 제조, 물류, 서비스 로봇 분야에서 multimodal 이해와 정밀한 제어가 동시에 필요한 상황에서 유용할 것으로 기대된다.