RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, Stanley T. Birchfield

arXiv:2411.16537 · 2026-07-27 공개 · arXiv · PDF

vision-language robot-manipulation robotics spatial-reasoning dataset spatial-affordance spatial-understanding egocentric-images

Abstract

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant challenges in spatial reasoning tasks, as their training data are based on general-purpose image datasets that often lack sophisticated spatial understanding. For example, datasets frequently do not capture reference frame comprehension, yet effective spatial reasoning requires understanding whether to reason from ego-, world-, or object-centric perspectives. To address this issue, we introduce RoboSpatial, a large-scale dataset for spatial understanding in robotics. It consists of real indoor and tabletop scenes, captured as 3D scans and egocentric images, and annotated with rich spatial information relevant to robotics. The dataset includes 1M images, 5k 3D scans, and 3M annotated spatial relationships, and the pairing of 2D egocentric images with 3D scans makes it both 2D- and 3D- ready. Our experiments show that models trained with RoboSpatial outperform baselines on downstream tasks such as spatial affordance prediction, spatial relationship prediction, and robotics manipulation.

한국어 요약

한 줄 요약

RoboSpatial은 2D 및 3D 시각-언어 모델의 공간 이해 능력을 향상시키기 위해 설계된 대규모 로봇용 데이터셋이다.

핵심 기여도

핵심 아이디어

기존 시각-언어 모델(VLM)은 일반 이미지 데이터셋에서 학습되어 공간 참조 프레임(ego-, world-, object-centric)을 이해하는 데 어려움을 겪는다. 이는 로봇이 실제 환경에서 의미 있는 상호작용을 하기 위해 필수적인 능력이다. RoboSpatial은 이러한 문제를 해결하기 위해 설계된 데이터셋으로, 공간 관계를 3가지 유형(공간 맥락, 공간 호환성, 공간 구성)으로 구분하고, 각 질문-답변 쌍을 3가지 참조 프레임에서 제시한다. 이는 모델이 다양한 관점에서 공간 정보를 해석하도록 유도하며, 실제 로봇 작업에 필요한 유연한 공간 추론 능력을 향상시킨다.

기술적 접근법

주요 결과

의의 및 한계

RoboSpatial은 로봇이 실제 환경에서 공간 정보를 해석하고 조작할 수 있도록 기반을 제공하며, 기존 VLM의 공간 이해 능력 한계를 보완한다. 특히, 2D 및 3D 모델 모두 사용 가능한 데이터셋 구조는 다양한 연구와 응용에 유리하다. 그러나 일부 3D VLM이 사용된 훈련 데이터와 유사한 환경에 노출되어 있어 성능 편향이 발생할 수 있다는 한계가 있다. 또한, 데이터셋의 자동 생성 파이프라인은 새로운 공간 관계나 환경에 확장 가능하지만, 수작업 검증이 필요한 부분도 존재한다.

실용적 활용

RoboSpatial은 로봇이 물체 배치, 경로 계획, 조작 등 실제 작업에 필요한 공간 이해 능력을 향상시키는 데 활용될 수 있다. 또한, 증강현실(AR) 및 자율주행 분야에서도 공간 관계 해석에 필요한 VLM 훈련에 사용될 수 있다.