VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Zhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu

arXiv:2609.19554 · 2026-09-20 공개 · arXiv · PDF

llm-evaluation long-horizon spatial-reasoning embodied-intelligence task-success active-perception visual-demonstrations metric-control

Abstract

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.

한국어 요약

한 줄 요약

VABench는 시각 시연, 능동적 관찰, 메트릭 제어를 통해 MLLM의 공간 지능을 종합적으로 평가하는 벤치마크이다.

핵심 기여도

핵심 아이디어

VABench는 기존의 고정/다중 뷰 추론에서 벗어나, 실제 물리적 상호작용 환경에서 모델이 능동적으로 증거를 수집하고, 공간적 판단을 메트릭 명령으로 변환하는 능력을 평가한다. 이는 단순히 시각 정보를 기반으로 객체 위치를 파악하는 것을 넘어, **미완전한 관찰 상황에서 필요한 정보를 식별하고, 공통 공간 프레임에서 해석한 후, 실행 가능한 행동으로 전환**하는 능력을 측정한다.

모델은 **RGB 시연만을 기반으로 학습**하며, **특정 객체 포즈, 오라클 경로, 학습된 액션 헤드 없이** 고정된 컨트롤러를 통해 명령을 실행한다. 이는 모델이 **시각적 시연을 기반으로 절차적 맥락을 이해하고, 그에 따라 메트릭 명령을 생성**해야 한다는 점에서 기존 접근과 구별된다.

기술적 접근법

주요 결과

의의 및 한계

VABench는 MLLM이 시각 시연을 기반으로 능동적 증거 수집과 메트릭 명령 생성을 수행할 수 있는지 종합적으로 평가하는 첫 번째 벤치마크로, **공간 지능의 실행 가능성과 일반화 능력을 측정**하는 데 기여한다. 특히, **공간 진단과 실행 성능 간의 불일치**, **기하학적 전이 시의 성능 저하**, **장기적 작업의 합성 실패**라는 세 가지 한계를 명확히 밝혀내며, MLLM의 실제 물리적 환경 적용 가능성에 대한 중요한 통찰을 제공한다.

하지만, **모든 실험은 시뮬레이션 환경에서 진행**되었으며, **실제 로봇 환경에서의 성능 검증이 필요**하다는 한계가 있다. 또한, **모델이 공간 판단을 메트릭 명령으로 변환하는 능력은 여전히 제한적**이며, **온라인 수정 능력도 약한 수준**으로 나타났다.

실용적 활용

VABench는 **로봇 조작, 물리적 환경에서의 자율 시스템 개발, MLLM의 실제 적용 가능성 평가**에 활용될 수 있다. 특히, **능동적 관찰과 메트릭 명령 생성 능력을 갖춘 MLLM의 개발 및 평가**에 중요한 기준이 될 수 있으며, **공간 지능 기반의 서비스 로봇, 산업 자동화 시스템** 등에 적용 가능하다.