VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Shiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu, Yuan Li, Qinghui Gao, Zhaoye Fei, Zhangyue Yin, Zuxuan Wu, Yu-Gang Jiang, Xipeng Qiu

arXiv:2412.18194 · 2026-07-27 공개 · arXiv · PDF

foundation-models vision-language-action spatial-reasoning vla long-horizon-reasoning task-generalization physical-laws robotics-benchmark

Abstract

General-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on foundation models especially Vision-Language-Action models (VLAs) have shown a substantial potential to solve language-conditioned manipulation (LCM) tasks well. However, existing benchmarks do not adequately meet the needs of VLAs and relative algorithms. To better define such general-purpose tasks in the context of LLMs and advance the research in VLAs, we present VLABench, an open-source benchmark for evaluating universal LCM task learning. VLABench provides 100 carefully designed categories of tasks, with strong randomization in each category of task and a total of $2000+$ objects. VLABench stands out from previous benchmarks in four key aspects: 1) tasks requiring world knowledge and common sense transfer, 2) natural language instructions with implicit human intentions rather than templates, 3) long-horizon tasks demanding multi-step reasoning, and 4) evaluation of both action policies and language model capabilities. The benchmark assesses multiple competencies including understanding of mesh&texture, spatial relationship, semantic instruction, physical laws, knowledge transfer and reasoning, etc. To support the downstream finetuning, we provide high-quality training data collected via an automated framework incorporating heuristic skills and prior information. The experimental results indicate that both the current state-of-the-art pretrained VLAs and the workflow based on VLMs face challenges in our tasks.11Codes and more videos are available at https://vlabench.github.io/22Corresponding to: [email protected], [email protected].

한국어 요약

한 줄 요약

VLABench는 장기적 계획과 다차원적 추론이 필요한 언어 조건화 로봇 조작을 평가하기 위한 대규모 벤치마크이다.

핵심 기여도

핵심 아이디어

VLABench는 기존의 로봇 조작 벤치마크가 Foundation Models, 특히 Vision-Language-Action (VLA) 모델의 요구를 충족하지 못한다는 점에서 출발하였다. 기존 태스크는 템플릿 기반 언어 지시와 단기적 행동에 초점을 맞추고 있었으며, 실제 세계에서 필요한 세계 지식, 추론, 장기적 계획 능력을 평가하지 못하였다. VLABench는 6가지 핵심 능력 (공감각 이해, 물리 법칙, 추론, 시공간 관계 등)을 평가하기 위해 100개의 태스크 범주를 정의하였다. 이 태스크는 자연스러운 언어 지시와 암묵적인 인간 의도를 반영하며, 다단계 추론이 필요한 장기적 태스크를 포함한다. 예를 들어, "테일러 스위프트에게 코카콜라 캔을 이동하라"는 태스크는 세계 지식과 지식 전이 능력을 요구하며, "포어오버 커피를 만드라"는 태스크는 복잡한 계획 능력을 평가한다.

기술적 접근법

VLABench는 6개의 핵심 평가 영역 (공감각 이해, 물리 법칙, 추론, 시공간 관계 등)을 기준으로 100개의 태스크 범주를 정의하였다. 태스크는 2000개 이상의 3D 객체와 다양한 장면을 활용하여 강한 랜덤화를 적용하였다. 평가 프로세스는 다음과 같다:

주요 결과

의의 및 한계

VLABench는 로봇 조작 분야에서 Foundation Models의 능력을 종합적으로 평가하는 첫 번째 벤치마크로, 학술적·실용적 가치가 크다. 특히, 세계 지식, 추론, 장기적 계획 등 실제 세계에서 필요한 능력을 평가함으로써, 기존의 단기적 행동 중심의 평가를 넘어선다. 그러나 VLABench는 시뮬레이션 기반 평가이므로, 실제 로봇 환경에서의 성능과의 차이가 있을 수 있다. 또한, 현재 VLAs와 VLMs는 VLABench의 태스크에서 제한적인 성능을 보여, 로봇 스케일링 연구에 대한 불확실성이 남아 있다.

실용적 활용

VLABench는 로봇 조작 알고리즘 개발자들이 Foundation Models의 능력을 종합적으로 평가하고, 장기적 추론과 물리 법칙 이해를 향상시키는 데 활용될 수 있다. 특히, VLM 기반 워크플로우나 VLA 사전 학습 모델의 미세조정 및 성능 개선에 유용하며, 실제 산업 환경에서의 로봇 시스템 설계에도 기초 자료로 활용될 수 있다.