vlm benchmarking image-generation visual-reasoning generative-models spatial-reasoning pixel-space spatial-cognition
Abstract
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.
한국어 요약
한 줄 요약
ProVisE와 SpatialGen-Bench를 통해 이미지 생성 모델의 시각적 공간 인지 평가가 가능해졌다.
핵심 기여도
- ProVisE: 이미지 생성 모델의 시각적 답변을 원래 평가 지표와 호환되는 구조화된 예측으로 변환하는 프레임워크.
- Agentic builder: 새로운 벤치마크에 맞는 프로토콜을 자동으로 생성하고 검증하는 기능 포함.
- SpatialGen-Bench: 14개의 공간 하위 과제, 4개의 능력 수준, 다양한 답변 형태를 갖춘 470개 샘플의 진단 벤치마크.
- 이미지 생성 모델이 픽셀 공간에서 직접 공간 판단을 표현할 때 경쟁력 있음을 실험적으로 입증.
핵심 아이디어
기존 공간 추론 벤치마크는 텍스트 또는 좌표 기반 답변을 요구하므로, 이미지 생성 모델의 시각적 답변을 평가하기 어려웠다. 이 연구는 ProVisE라는 프레임워크를 통해, 이미지 생성 모델이 프로토콜에 따라 시각적 답변을 생성하게 한 후 이를 구조화된 예측으로 변환하여 기존 평가 지표와 호환되게 만든다. Agentic builder는 새로운 벤치마크에 맞는 프로토콜을 자동 생성하여, 평가 프로세스를 확장 가능하게 한다. 이는 텍스트 출력 VLM과 이미지 생성 모델을 동일한 공간 과제 의미하에서 비교할 수 있는 기반을 제공한다.
기술적 접근법
- **ProVisE**: 이미지 생성 모델이 프로토콜에 따라 시각적 답변을 생성하도록 유도하고, 이를 구조화된 예측으로 파싱하여 기존 평가 지표와 호환되게 만든다.
- **Agentic builder**: 새로운 벤치마크를 작업 수준 계약으로 정규화하고, 각 작업에 적합한 생성-파서 프로토콜을 생성 및 검증.
- **SpatialGen-Bench**: 14개의 공간 하위 과제, 4개의 능력 수준, 다양한 답변 형태를 포함한 470개 샘플로 구성된 진단 벤치마크.
- **평가 대상 모델**: 대표적인 텍스트 출력 VLM과 이미지 생성 모델 (예: FLUX, Seedream, Janus)을 동일한 설정에서 평가.
- **외부 벤치마크 검증**: ProVisE의 프로토콜 전이 능력을 6개 외부 공간 벤치마크에서 검증.
주요 결과
- 이미지 생성 모델은 공간 답변이 픽셀 공간에서 직접 표현될 수 있을 때 경쟁력 있음.
- 텍스트 출력 VLM은 구성적 공간 추론에서 명확한 우위를 보임.
- 실험 결과, ProVisE는 외부 공간 벤치마크에서도 프로토콜 전이가 가능함을 입증.
- SpatialGen-Bench에서 실패한 시각적 답변은 대부분 파서 가능하지만 잘못된 답변을 포함함.
의의 및 한계
ProVisE와 SpatialGen-Bench는 텍스트 출력 VLM과 이미지 생성 모델을 동일한 공간 과제 의미하에서 비교할 수 있는 첫 번째 메트릭 호환 평가베이스를 제공한다. 이는 시각적 표현과 텍스트 기반 추론의 보완적 강점을 밝히는 데 기여한다. 그러나, 이미지 생성 모델은 구성적 추론에서 여전히 약점을 보이며, 시각적 표현이 추론을 대체하지 못함을 보여준다. 또한, 파서의 신뢰도가 최종 평가 결과에 영향을 줄 수 있다는 한계도 존재한다.
실용적 활용
ProVisE와 SpatialGen-Bench는 로봇, 자율 주행, AR/VR 등 시각적 공간 인지가 필요한 산업 분야에서 모델 평가에 활용될 수 있다. 또한, 텍스트 기반 추론과 시각적 표현의 결합이 필요한 연구 분야에서 이중 표현 시스템 설계에 기초 자료로 활용 가능하다.