HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Zhaorun Chen, Zhaorun Chen, Zhuokai Zhao, Hongyin Luo, Huaxiu Yao, Bo Li, Jiawei Zhou

arXiv:2403.00425 · 2026-07-27 공개 · arXiv · PDF

vision-language beam-search context-aware plug-and-play large-vision-language-models object-hallucination auto-focal-grounding decoding-algorithm

Abstract

While large vision-language models (LVLMs) have demonstrated impressive capabilities in interpreting multi-modal contexts, they invariably suffer from object hallucinations (OH). We introduce HALC, a novel decoding algorithm designed to mitigate OH in LVLMs. HALC leverages distinct fine-grained optimal visual information in vision-language tasks and operates on both local and global contexts simultaneously. Specifically, HALC integrates a robust auto-focal grounding mechanism (locally) to correct hallucinated tokens on the fly, and a specialized beam search algorithm (globally) to significantly reduce OH while preserving text generation quality. Additionally, HALC can be integrated into any LVLMs as a plug-and-play module without extra training. Extensive experimental studies demonstrate the effectiveness of HALC in reducing OH, outperforming state-of-the-arts across four benchmarks.

한국어 요약

한 줄 요약

HALC는 LVLM에서 발생하는 오브젝트 환각(OH)을 줄이기 위해 설계된 플러그 앤 플레이 디코딩 알고리즘이다.

핵심 기여도

핵심 아이디어

HALC는 기존 LVLM에서 흔히 발생하는 OH 문제를 해결하기 위해 **로컬과 글로벌 수준의 동시 처리**를 강조한다. 로컬 수준에서는 이미 생성된 토큰이 환각인지 실시간으로 감지하고, **adaptive focal-contrast grounding** 메커니즘을 통해 시각 정보와 일치하는 토큰 확률을 재분배한다. 글로벌 수준에서는 **matching-based beam search**를 사용하여 시각 매칭 점수를 기반으로 최종 생성 텍스트를 조정함으로써 OH를 줄이면서 텍스트 생성 품질을 유지한다. 이는 기존의 OH 감소 방법들이 주로 존재 환각에만 집중하고, 속성 및 관계 환각은 자동 생성에 의존하는 한계를 극복한다.

기술적 접근법

주요 결과

의의 및 한계

HALC는 OH 감소와 텍스트 생성 품질을 동시에 유지하는 데 성공했으며, 기존 방법들이 외부 모델이나 추가 데이터에 의존하는 문제를 해결한 점에서 학술적·실용적 가치가 높다. 또한, 플러그 앤 플레이 구조로 기존 LVLM에 쉽게 적용 가능하다는 점이 혁신적이다. 그러나 HALC는 **속성 및 관계 환각 감소의 정량적 평가**는 명시되지 않았으며, **추가 학습 없이 적용 가능한 범용성**에도 한계가 있을 수 있다.

실용적 활용

HALC는 이미지 캡셔닝, 시각 질의 응답(VQA), 멀티모달 대화 시스템 등에서 오브젝트 환각을 줄이면서 텍스트 생성 품질을 유지해야 하는 산업 및 연구 상황에 적용 가능하다. 특히, LVLM 기반의 챗봇, 자동 번역, 멀티모달 검색 시스템 등에 유용하게 활용될 수 있다.