MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen

arXiv:2609.35652 · 2026-10-06 공개 · arXiv · PDF

vision-language-model mobile-manipulation robocasa365 ebench masked-attention robot-pretraining arm-base-coordination future-branch

Abstract

Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.

한국어 요약

한 줄 요약

MM-ABC는 이동 및 조작의 협업을 학습하는 모바일 조작 기반 모델로, 다양한 데이터에서 61.2% 성공률을 달성.

핵심 기여도

핵심 아이디어

기존 모바일 조작 연구는 3D 재구성이나 예측 모델을 사용해 공간 인식을 강화했으나, 이는 계산 비용이 높고, 이동과 조작 스트림 간 협업이 부족했다. MM-ABC는 이동과 조작을 별도 스트림으로 유지하면서도 **masked joint attention**을 통해 정보를 공유하는 **MM-APT**를 도입했다. 또한, **clean-action x-prediction**을 사용해 정확한 엔드포인트 예측을 가능하게 하며, **future branch**를 통해 기하학적 의도를 학습함으로써 훈련 신호를 강화했다. 이는 **velocity prediction** 대비 더 정확한 작업 분배와 노이즈 제거를 가능하게 한다는 점에서 차별화된다.

기술적 접근법

주요 결과

의의 및 한계

MM-ABC는 이동과 조작을 별도 스트림으로 유지하면서도 협업을 가능하게 하며, 이종 로봇 데이터를 기반으로 사전 훈련하여 다양한 환경에서 성능을 발휘한다. 특히, **clean-action x-prediction**과 **future branch**를 통해 훈련 신호를 강화하고, 추론 비용을 절감하는 점에서 실용적 가치가 있다. 그러나 **MM-ABC**는 특정 작업에 대한 미세 조정 없이도 성능을 유지하지만, 복잡한 환경 변화에 대한 일반화 능력은 추가 연구가 필요하다. 또한, **velocity prediction** 대비 **clean-action x-prediction**의 장점을 명확히 입증했으나, 이는 특정 데이터셋에서만 검증된 점이 한계다.

실용적 활용

MM-ABC는 로봇이 다양한 환경에서 이동과 조작을 동시에 수행해야 하는 상황, 예를 들어 가정용 서비스 로봇, 물류 시스템, 산업 자동화 등에 적용 가능하다. 특히, **MM-APT**와 **clean-action x-prediction**은 복잡한 작업 분배와 정확한 엔드포인트 예측이 필요한 실제 작업에서 유용하다.