Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang

arXiv:2607.21072 · 2026-07-24 공개 · arXiv · PDF

vlm benchmarking image-generation visual-reasoning generative-models spatial-reasoning pixel-space spatial-cognition

Abstract

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.

한국어 요약

한 줄 요약

ProVisE와 SpatialGen-Bench를 통해 이미지 생성 모델의 시각적 공간 인지 평가가 가능해졌다.

핵심 기여도

핵심 아이디어

기존 공간 추론 벤치마크는 텍스트 또는 좌표 기반 답변을 요구하므로, 이미지 생성 모델의 시각적 답변을 평가하기 어려웠다. 이 연구는 ProVisE라는 프레임워크를 통해, 이미지 생성 모델이 프로토콜에 따라 시각적 답변을 생성하게 한 후 이를 구조화된 예측으로 변환하여 기존 평가 지표와 호환되게 만든다. Agentic builder는 새로운 벤치마크에 맞는 프로토콜을 자동 생성하여, 평가 프로세스를 확장 가능하게 한다. 이는 텍스트 출력 VLM과 이미지 생성 모델을 동일한 공간 과제 의미하에서 비교할 수 있는 기반을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

ProVisE와 SpatialGen-Bench는 텍스트 출력 VLM과 이미지 생성 모델을 동일한 공간 과제 의미하에서 비교할 수 있는 첫 번째 메트릭 호환 평가베이스를 제공한다. 이는 시각적 표현과 텍스트 기반 추론의 보완적 강점을 밝히는 데 기여한다. 그러나, 이미지 생성 모델은 구성적 추론에서 여전히 약점을 보이며, 시각적 표현이 추론을 대체하지 못함을 보여준다. 또한, 파서의 신뢰도가 최종 평가 결과에 영향을 줄 수 있다는 한계도 존재한다.

실용적 활용

ProVisE와 SpatialGen-Bench는 로봇, 자율 주행, AR/VR 등 시각적 공간 인지가 필요한 산업 분야에서 모델 평가에 활용될 수 있다. 또한, 텍스트 기반 추론과 시각적 표현의 결합이 필요한 연구 분야에서 이중 표현 시스템 설계에 기초 자료로 활용 가능하다.