vision-language inference-time hallucination-detection lvlm object-hallucination semantic-refinement visual-attention multi-head-attention
Abstract
Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspect rather than the visual. In this paper, we address how LVLMs process visual information and whether this process causes hallucination. Firstly, we use the attention lens to identify the stages at which LVLMs handle visual data, discovering that the middle layers are crucial. Moreover, we find that these layers can be further divided into two stages: "visual information enrichment" and "semantic refinement" which respectively propagate visual data to object tokens and interpret it through text. By analyzing attention patterns during the visual information enrichment stage, we find that real tokens consistently receive higher attention weights than hallucinated ones, serving as a strong indicator of hallucination. Further examination of multi-head attention maps reveals that hallucination tokens often result from heads interacting with inconsistent objects. Based on these insights, we propose a simple inference-time method that adjusts visual attention by integrating information across various heads. Extensive experiments demonstrate that this approach effectively mitigates hallucinations in mainstream LVLMs without additional training costs.1
한국어 요약
한 줄 요약
LVLM에서 시각 정보 처리 단계를 분석하여 오브젝트 환영을 감지하고 완화하는 새로운 방법을 제안한다.
핵심 기여도
- 중간 레이어를 두 단계로 분류: "visual information enrichment"와 "semantic refinement".
- 실제 토큰이 환영 토큰보다 더 높은 attention weight를 갖는다는 점을 발견.
- Multi-Head Self-Attention (MHSA)에서 불일치 객체와 상호작용하는 헤드가 환영을 유발한다는 통찰.
- 추론 단계에서 헤드 간 정보 통합을 통해 6.3~24.1 포인트의 환영 감소를 달성.
핵심 아이디어
LVLM에서 오브젝트 환영이 발생하는 주요 원인은 시각 정보 처리 단계, 특히 중간 레이어에서의 불확실한 객체 상호작용이다. 기존 연구는 언어 편향에 집중했으나, 본 연구는 시각 정보 처리 과정을 분석함으로써 새로운 통찰을 제시한다. "Visual Attention Ratio (VAR)"를 도입하여 시각 정보의 레이어별 분포를 분석하고, 중간 레이어가 시각 정보 처리에 핵심적인 역할을 한다는 것을 밝혔다. 특히, "visual information enrichment" 단계에서 실제 객체 토큰이 더 높은 attention weight를 가지며, MHSA 헤드 간 불일치 객체 상호작용이 환영을 유발한다는 점이 핵심 발견이다.
기술적 접근법
- **Visual Attention Ratio (VAR)**: 시각 정보의 레이어별 attention 분포를 측정하는 지표.
- **Logit Lens**: 이미지 토큰의 hidden state를 LVLM의 어휘로 디코딩하여 시각 정보 흐름 분석.
- **Multi-Head Self-Attention (MHSA)**: 헤드 간 attention map 시각화를 통해 불일치 객체 상호작용 분석.
- **Inference-Time Method**: 헤드 간 정보 통합을 통해 attention 조정. 추가 학습 없이 적용 가능.
주요 결과
- LLaVA-1.5-7B 데이터셋에서 AUROC 74%, mAP 88% 달성.
- CHAIR I 지표에서 6.3 포인트, CHAIR S 지표에서 24.1 포인트의 환영 감소.
- 기존 모델 대비, 추론 단계에서의 오브젝트 환영 감소 효과가 입증됨.
의의 및 한계
- LVLM의 환영 메커니즘을 시각 정보 처리 관점에서 체계적으로 분석한 최초의 연구.
- 추론 단계에서의 간단한 attention 조정으로 환영을 완화할 수 있다는 실용적 가치.
- 그러나 본 연구는 특정 LVLM 아키텍처에만 적용되었으며, 다양한 모델 간 일반화 가능성은 추가 연구 필요.
실용적 활용
- 이미지 캡셔닝, 시각 질의 응답 등 LVLM이 활용되는 산업 분야에서 환영 감소에 활용 가능.
- 추론 단계에서의 실시간 attention 조정을 통해 모델 신뢰도를 향상시킬 수 있음.
- 추후 연구에서는 토큰 분류(예: 색상, 텍스트, 객체)를 기반으로 보다 세분화된 시각 정보 처리 분석이 기대됨.