Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, Hongsheng Li

arXiv:2403.16999 · 2026-07-27 공개 · arXiv · PDF

chain-of-thought mllm multi-modal vqa large-scale-dataset interpretability bounding-box visual-cot

Abstract

Multi-Modal Large Language Models (MLLMs) have demonstrated impressive performance in various VQA tasks. However, they often lack interpretability and struggle with complex visual inputs, especially when the resolution of the input image is high or when the interested region that could provide key information for answering the question is small. To address these challenges, we collect and introduce the large-scale Visual CoT dataset comprising 438k question-answer pairs, annotated with intermediate bounding boxes highlighting key regions essential for answering the questions. Additionally, about 98k pairs of them are annotated with detailed reasoning steps. Importantly, we propose a multi-turn processing pipeline that dynamically focuses on visual inputs and provides interpretable thoughts. We also introduce the related benchmark to evaluate the MLLMs in scenarios requiring specific local region identification. Extensive experiments demonstrate the effectiveness of our framework and shed light on better inference strategies. The Visual CoT dataset, benchmark, and pre-trained models are available on https://hao-shao.com/projects/viscot.html to support further research in this area.

한국어 요약

한 줄 요약

Visual CoT는 438k 질문-답변 쌍을 포함한 대규모 데이터셋과 MLLM의 해석 가능성 향상을 위한 다단계 처리 파이프라인을 제안한다.

핵심 기여도

핵심 아이디어

기존 MLLM은 시각 정보 처리 시 해석 가능성과 정확도에 한계가 있으며, 특히 고해상도 이미지나 작은 관심 영역에 대한 처리가 어려움. 이를 해결하기 위해 Visual CoT는 질문-답변 쌍에 바운딩 박스와 추론 단계를 추가하여 모델이 핵심 영역을 동적으로 식별하도록 유도함. 이는 인간이 시각 정보를 처리하는 방식을 모방한 것으로, 모델이 전체 이미지를 스캔한 후 핵심 영역에 집중하는 다단계 추론 과정을 구현함. 특히, Visual CoT 파이프라인은 원본 이미지와 세부 지역 이미지를 모두 이해하여 최종 답변을 생성함.

기술적 접근법

주요 결과

의의 및 한계

Visual CoT는 MLLM의 해석 가능성과 시각 추론 능력을 동시에 향상시키는 새로운 접근법을 제시함. 특히, 바운딩 박스와 추론 단계를 통해 모델이 어떻게 답변에 도달하는지를 명확히 보여주며, 연구자들이 추론 과정을 분석하고 개선할 수 있도록 지원함. 그러나, 현재 데이터셋은 5개 도메인에만 제한되어 있으며, 더 다양한 시각적 상황을 다루는 데이터 확장을 필요로 함. 또한, 바운딩 박스 예측 정확도가 모델 전체 성능에 큰 영향을 미치므로, 정확한 객체 식별 기술의 개선이 필수적임.

실용적 활용

Visual CoT는 의료 영상 분석, 자율 주행, 보안 감시 등에서 시각 정보를 정확히 해석하고 추론하는 MLLM의 활용을 촉진할 수 있음. 특히, 시각 추론 과정을 해석 가능한 형태로 제공함으로써, 모델 신뢰도를 높이고 오류 원인을 분석하는 데 유용함.