MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao, Tong Zhang, Huan Zhang

arXiv:2609.38078 · 2026-10-05 공개 · arXiv · PDF

vlm vision-language-action robot-control visual-grounding libero-pro embodied-reasoning asynchronous-execution zero-shot-manipulation

Abstract

Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.

한국어 요약

한 줄 요약

MotorMind는 VLM 기반의 Zero-Shot 로봇 조작 시스템으로, LIBERO-PRO에서 66.7% 성공률을 달성하며 xArm6 로봇에서도 95% 성공률을 보인다.

핵심 기여도

핵심 아이디어

기존 VLA 모델은 Zero-Shot 조작에서 제한적이며, 외부 도구에 의존하는 경향이 있다. 반면, MotorMind는 **VLM 자체가 외부 도구 없이 로봇을 조작**할 수 있도록 설계되었다. 핵심 아이디어는 **VLM이 중간 수준 행동(mid-level action)**을 제안하고, 이를 결정론적 제어 시스템과 연결하는 것이다. 이는 인간 텔레오퍼레이터처럼 **관측을 기반으로 행동을 직접 내리고, 실행 피드백에 따라 조정**할 수 있도록 한다. 예를 들어, **Planner, Executor, Monitor, Verifier, Memory**라는 5가지 역할을 갖는 VLM은 각각 다른 맥락과 출력 권한을 가지며, **Controller**가 물리적 실행을 담당한다.

기술적 접근법

주요 결과

의의 및 한계

MotorMind는 **VLM을 외부 도구 없이 직접 로봇 조작에 활용**할 수 있음을 보여준다. 이는 Zero-Shot 로봇 조작을 **VLM의 일반적 능력과 제어 표현의 일치 문제로 재정의**하는 새로운 접근법이다. 그러나 실패 분석에서 **시각적 지정(visual grounding), 체화적 추론(embodied reasoning), 행동 지식(action knowledge)** 부족이 주요 원인으로 드러났다. 즉, VLM의 시각 이해와 자기 검증 능력이 향상될수록 이 시스템의 성능도 개선될 수 있다. 또한, **Planner 제거 시 성공률 0%**로 떨어지는 점에서, 계획 역할은 필수적임을 알 수 있다.

실용적 활용

MotorMind는 **실제 로봇(xArm6)에서 즉석 조작 및 인간 방해 상황에서도 높은 성공률**을 보이므로, **로봇 제조, 물류, 서비스 산업** 등에서 즉시 적용 가능하다. 또한, **강력한 VLM을 사용하면 추가 학습 없이도 성능 향상**이 가능하므로, **빠르게 발전하는 VLM 기반의 로봇 시스템 개발**에 기여할 수 있다.