BLINK: Multimodal Large Language Models Can See but Not Perceive

Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna

arXiv:2404.12390 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation multimodal-llms depth-estimation visual-perception computer-vision image-understanding visual-correspondence multi-view-reasoning

Abstract

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans"within a blink"(e.g., relative depth estimation, visual correspondence, forensics detection, and multi-view reasoning). However, we find these perception-demanding tasks cast significant challenges for current multimodal LLMs because they resist mediation through natural language. Blink reformats 14 classic computer vision tasks into 3,807 multiple-choice questions, paired with single or multiple images and visual prompting. While humans get 95.70% accuracy on average, Blink is surprisingly challenging for existing multimodal LLMs: even the best-performing GPT-4V and Gemini achieve accuracies of 51.26% and 45.72%, only 13.17% and 7.63% higher than random guessing, indicating that such perception abilities have not"emerged"yet in recent multimodal LLMs. Our analysis also highlights that specialist CV models could solve these problems much better, suggesting potential pathways for future improvements. We believe Blink will stimulate the community to help multimodal LLMs catch up with human-level visual perception.

한국어 요약

한 줄 요약

BLINK는 시각 인지 능력을 평가하는 새로운 멀티모달 LLM 벤치마크로, 인간은 쉽게 풀지만 기존 모델은 어려움을 겪는다.

핵심 기여도

핵심 아이디어

BLINK는 자연어로 매개되지 않는 핵심 시각 인지 능력을 평가하는 데 초점을 맞춘다. 기존 멀티모달 LLM 평가가 시맨틱 이해나 객체 인식에 집중했다면, BLINK는 상대 깊이 추정, 시각 대응, 포렌식 탐지, 다중 뷰 추론 등 인간이 직관적으로 처리하는 태스크를 포함한다. 이러한 작업들은 시각 정보를 단순히 텍스트로 변환하는 것 이상의 인지 능력을 요구하며, 이는 기존 멀티모달 LLM이 부족한 영역임을 드러낸다.

기술적 접근법

주요 결과

의의 및 한계

BLINK는 멀티모달 LLM의 시각 인지 능력 부족을 명확히 드러내며, 시각 정보 처리 기술 발전을 촉진할 수 있는 새로운 평가 기준을 제시한다. 그러나 BLINK는 이미지 기반 질문에만 초점을 맞추고 있어, 동영상, 3D 데이터 등 다양한 시각 입력을 다루는 평가가 필요하다는 한계가 있다. 또한, 멀티모달 LLM이 시각 인지 능력을 어떻게 발현시킬 수 있을지에 대한 구체적 해결 방안은 제시되지 않았다.

실용적 활용

BLINK는 멀티모달 LLM의 시각 인지 능력을 개선하려는 연구자들에게 유용한 평가 도구로 활용될 수 있으며, 특히 자율 주행, 의료 영상 분석, 보안 감시 등 시각 인지가 중요한 산업 분야에서 모델 성능 검증에 사용될 수 있다.