Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

Zhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo, Yankun Shen, Xu Yang

arXiv:2411.16724 · 2026-07-27 공개 · arXiv · PDF

vision-language inference-time hallucination-detection lvlm object-hallucination semantic-refinement visual-attention multi-head-attention

Abstract

Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspect rather than the visual. In this paper, we address how LVLMs process visual information and whether this process causes hallucination. Firstly, we use the attention lens to identify the stages at which LVLMs handle visual data, discovering that the middle layers are crucial. Moreover, we find that these layers can be further divided into two stages: "visual information enrichment" and "semantic refinement" which respectively propagate visual data to object tokens and interpret it through text. By analyzing attention patterns during the visual information enrichment stage, we find that real tokens consistently receive higher attention weights than hallucinated ones, serving as a strong indicator of hallucination. Further examination of multi-head attention maps reveals that hallucination tokens often result from heads interacting with inconsistent objects. Based on these insights, we propose a simple inference-time method that adjusts visual attention by integrating information across various heads. Extensive experiments demonstrate that this approach effectively mitigates hallucinations in mainstream LVLMs without additional training costs.1

한국어 요약

한 줄 요약

LVLM에서 시각 정보 처리 단계를 분석하여 오브젝트 환영을 감지하고 완화하는 새로운 방법을 제안한다.

핵심 기여도

핵심 아이디어

LVLM에서 오브젝트 환영이 발생하는 주요 원인은 시각 정보 처리 단계, 특히 중간 레이어에서의 불확실한 객체 상호작용이다. 기존 연구는 언어 편향에 집중했으나, 본 연구는 시각 정보 처리 과정을 분석함으로써 새로운 통찰을 제시한다. "Visual Attention Ratio (VAR)"를 도입하여 시각 정보의 레이어별 분포를 분석하고, 중간 레이어가 시각 정보 처리에 핵심적인 역할을 한다는 것을 밝혔다. 특히, "visual information enrichment" 단계에서 실제 객체 토큰이 더 높은 attention weight를 가지며, MHSA 헤드 간 불일치 객체 상호작용이 환영을 유발한다는 점이 핵심 발견이다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용