Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Hanoona Rasheed, Mohammed Irfan Kurpath, Bin Ren, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan

arXiv:2609.35718 · 2026-09-29 공개 · arXiv · PDF

visual-reasoning reconstruction computer-vision temporal-consistency gpt-6-astra general-purpose-ai structured-prediction semantic-interpretation

Abstract

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We evaluate GPT-6 Astra alongside five frontier general-purpose AI systems across 34 capabilities and 55 benchmarks spanning nine areas of computer vision. We compare their performance with dedicated models and humans where suitable references are available. Astra demonstrates broad visual capability, with substantial gains over other frontier systems in visual and spatial reasoning and several forms of structured prediction. Across the state-of-the-art systems, a consistent pattern emerges. Capabilities involving semantic interpretation, reasoning, and object-centric prediction increasingly approach or reach available reference levels. In contrast, larger gaps remain when tasks require metric geometric accuracy, faithful reconstruction, temporally consistent dense prediction, or specialized fine-grained visual knowledge. Additional reasoning and specialist tools close selected gaps, but their benefits vary across capabilities. These results map a changing landscape of computer vision in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.

한국어 요약

한 줄 요약

GPT-6 Astra는 시각 추론과 구조화된 예측에서 기존 최첨단 시스템을 앞선다.

핵심 기여도

핵심 아이디어

전통적으로 전용 모델이 담당했던 다양한 컴퓨터 비전 작업이 일반 목적 시스템으로 확장되고 있다. 이 연구는 GPT-6 Astra와 5개의 최첨단 일반 목적 AI 시스템을 34개 능력, 55개 벤치마크에서 평가함으로써, 이러한 시스템이 어디까지 확장되었는지, 어떤 작업이 여전히 어려운지 분석했다. 핵심 통찰은 시각 정보가 의미 해석과 추론을 지원할 때, 일반 목적 시스템이 전문 모델 수준에 도달할 수 있다는 점이다. 반면, 정밀 기하학적 정확도, 신뢰성 있는 재구성, 시간적 일관성, 세부 시각 지식이 필요한 작업에서는 여전히 큰 격차가 남아 있다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 일반 목적 시스템이 점점 더 복잡한 시각 작업을 처리할 수 있음을 보여주며, 전용 모델의 역할이 변화하고 있음을 시사한다. 특히, 의미 해석과 추론을 기반으로 한 작업에서는 일반 목적 시스템이 전문 모델 수준에 도달할 수 있다. 그러나 정밀도와 신뢰성, 시간적 일관성 등이 필요한 작업에서는 여전히 전용 모델이 필요하다는 한계가 드러난다. 이는 컴퓨터 비전 연구가 정밀 시각 인식 분야에 집중해야 함을 시사한다.

실용적 활용

이 연구는 일반 목적 시스템이 객체 중심 예측, 시각 추론, 구조화된 예측 등에서 활용될 수 있음을 보여준다. 특히, 자연어 인터페이스를 통해 다양한 시각 작업을 수행하는 산업 분야(예: 로봇, 의료 영상 분석)에서 활용 가능하다. 그러나 정밀 기하학, 영상 재구성 등에서는 전용 모델과 혼합 사용이 필요할 수 있다.