RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan, Juekai Lin, Liang Liang, Zhuoyi Huang, Wenqiao Zhang, Juncheng Li, Siliang Tang, Jun Xiao, Yueting Zhuang
arXiv:2609.05324 · 2026-09-09 공개 · arXiv · PDF
vision-language-action robotics spatial-reasoning robotic-manipulation vla-models long-horizon-planning trajectory-data embodied-reasoning
Abstract
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.
한국어 요약
한 줄 요약
RoboSPA는 VLA 모델의 공간적·절차적 추론 능력을 평가하기 위한 대규모 로봇 조작 벤치마크이다.
핵심 기여도
- RoboSPA 데이터셋: 10개 태스크 범주, 56개 기본 태스크, 5단계 난이도로 280개 변형 태스크 제공.
- 527,000개의 트래젝토리 수집, 다양한 로봇 형태와 장면에서의 데이터 확보.
- 이진 성공률 외에도 진단적 평가 지표 도입.
- 대표적 VLA 모델 실험 결과: 복잡한 공간 관계, 정밀 실행, 메모리 집약적 계획에서 저조한 성능.
핵심 아이디어
기존 VLA 모델 평가가 단순한 장면과 짧은 수평 태스크에 제한되어 있어, 실제 환경에서의 추론 능력을 진단하기 어렵다는 문제를 지적한다. RoboSPA는 Fine-Grained Spatial Reasoning과 Long-Horizon Procedural Planning이라는 두 핵심 차원을 중심으로, 점점 복잡해지는 공간적 모호성과 절차적 복잡성을 반영한 태스크를 설계했다. 이는 로봇이 실제 세계에서의 공간 인식과 장기적 계획 능력을 평가하는 데 기여한다.
기술적 접근법
- RoboSPA는 56개의 기본 태스크를 5단계 난이도로 확장하여 280개의 태스크 변형을 구성.
- 527,000개의 트래젝토리 수집, 다양한 로봇 형태와 장면에서의 데이터 확보.
- 이진 성공률 외에도, 정밀도, 기억력, 절차적 일관성 등 진단적 평가 지표를 도입.
- 대표적 VLA 모델 실험을 통해 복잡한 공간 관계와 메모리 집약적 계획에서의 성능 저하를 분석.
주요 결과
- 대표적 VLA 모델은 복잡한 공간 관계와 정밀 실행에서 성능 저하.
- Long-Horizon Procedural Planning 태스크에서 실패율 증가, 메모리 집약적 계획에서 낮은 정확도.
- RoboSPA는 기존 VLA 모델의 한계를 명확히 진단하는 벤치마크로 제시됨.
의의 및 한계
RoboSPA는 VLA 모델의 실제 세계 적용 가능성과 추론 능력을 평가하는 데 중요한 진단 도구로 기여한다. 특히, 공간적·절차적 복잡성이 증가할수록 모델의 한계가 드러나는 점을 통해, 보다 강력하고 일반화된 로봇 에이전트 개발에 기초를 제공한다. 그러나 RoboSPA는 특정 로봇 형태와 환경에 기반한 데이터셋이므로, 보다 다양한 로봇 플랫폼과 장면에서의 확장이 필요하다.
실용적 활용
RoboSPA는 로봇 개발자와 연구자들이 VLA 모델의 실제 추론 능력을 평가하고, 복잡한 환경에서의 성능을 개선하는 데 활용할 수 있다. 특히, 장기적 계획과 정밀한 공간 인식이 필요한 산업 자동화, 서비스 로봇 분야에서 유용하게 사용될 수 있다.