An Exam for Active Observers

Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger

arXiv:2607.16165 · 2026-07-23 공개 · arXiv · PDF

vision-language code-generation mllm reasoning visual-perception perception-reasoning active-observation activevision

Abstract

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.

한국어 요약

한 줄 요약

ActiveVision 벤치마크를 통해 MLLMs가 활성 시각 관찰 능력을 갖추지 못함을 실증적으로 밝힘.

핵심 기여도

핵심 아이디어

인간의 시각은 단일 이미지가 아닌 가설에 따라 반복적으로 시선을 이동시키는 ‘활성 시각 관찰’로 이루어진다. 이 연구는 MLLMs가 이러한 능력을 갖추고 있는지, 기존 벤치마크가 이를 측정하지 못함을 지적하며, ActiveVision이라는 새로운 벤치마크를 제안한다. 이 벤치마크는 Distributed Scanning, Sequential Traversal, Visual Attribute Transfer 세 가지 인지 요구를 반영한 17개 태스크로 구성되며, 단일 이미지 설명으로는 해결 불가능한 구조적 복잡성을 가진다. GPT-image-2를 활용해 실제 이미지로 변환하면서 정확한 기하 구조를 유지하도록 설계되어, MLLMs가 반복적인 시각 인식을 요구하는 문제를 해결해야 성공할 수 있도록 한다.

기술적 접근법

주요 결과

의의 및 한계

ActiveVision은 MLLMs의 활성 시각 관찰 능력을 정량적으로 평가할 수 있는 첫 벤치마크로, 기존 벤치마크가 반복적 시각 인식을 측정하지 못하는 한계를 보완한다. 연구 결과는 MLLMs가 단일 시각 인코딩에 의존하며, 가설 기반 반복 시선 이동이 부족함을 보여준다. 그러나 ActiveVision은 특정 태스크 집합에 국한되며, 다양한 시각 상황을 포괄하지는 못한다. 또한, 툴 사용 시 일부 태스크에서 성능 향상이 있으나, 전체적으로는 여전히 인간 수준에 도달하지 못한다.

실용적 활용

ActiveVision은 로봇, 제조, 의료, 과학 탐구 등 반복적 시각 인식이 필요한 분야에서 MLLMs의 성능을 평가하는 기준으로 활용 가능하다. 활성 시각 관찰 능력을 갖춘 모델 개발은 실제 세계에서의 신뢰성 있는 AI 적용을 위한 필수 조건이다.