A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models

Pranab Sahoo, Prabhash Meharia, Akash Ghosh, Sriparna Saha, Vinija Jain, Aman Chadha

arXiv:2405.09589 · 2026-07-27 공개 · arXiv · PDF

large-language-models foundation-models image-generation hallucination-detection multimodal hallucination-mitigation taxonomy reliability

Abstract

The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks. However, the proliferation of FMs brings forth a critical challenge: the potential to generate hallucinated outputs, particularly in high-stakes applications. The tendency of foundation models to produce hallucinated content arguably represents the biggest hindrance to their widespread adoption in real-world scenarios, especially in domains where reliability and accuracy are paramount. This survey paper presents a comprehensive overview of recent developments that aim to identify and mitigate the problem of hallucination in FMs, spanning text, image, video, and audio modalities. By synthesizing recent advancements in detecting and mitigating hallucination across various modalities, the paper aims to provide valuable insights for researchers, developers, and practitioners. Essentially, it establishes a clear framework encompassing definition, taxonomy, and detection strategies for addressing hallucination in multimodal foundation models, laying the foundation for future research in this pivotal area.

한국어 요약

한 줄 요약

대규모 기초 모델(FMs)에서 발생하는 환각(hallucination) 문제를 다중 모달에서 체계적으로 분류하고 해결 방안을 제시한 종합적 서베이 논문.

핵심 기여도

핵심 아이디어

기초 모델(FMs)은 다양한 모달에서 뛰어난 성능을 보이지만, 생성된 결과가 실제 정보와 일치하지 않는 환각(hallucination)을 유발할 수 있다. 이는 특히 의료, 법률, 자동화 시스템 등 신뢰성과 정확성이 중요한 분야에서 큰 문제를 일으킬 수 있다. 본 논문은 환각을 단순히 오류로 보는 것이 아니라, 모델의 구조적 한계와 학습 데이터의 편향성, 추론 과정의 불확실성 등 다양한 요인으로 인해 발생한다고 분석한다. 이를 바탕으로 환각을 탐지하고 완화하기 위한 기술적 접근을 종합적으로 정리하고, 향후 연구 방향을 제시한다.

기술적 접근법

주요 결과

의의 및 한계

본 논문은 FMs에서 발생하는 환각 문제를 체계적으로 정리하고, 탐지 및 완화 기법을 종합적으로 분석함으로써, 연구자와 개발자에게 중요한 참고 자료를 제공한다. 특히, 다중 모달 환경에서의 환각 문제를 분류한 체계적 틀은 향후 연구의 기초가 될 수 있다. 그러나, 실제 환경에서의 환각 탐지 정확도는 여전히 낮은 수준이며, 모델별, 도메인별 차이가 크다는 점이 한계로 지적된다. 또한, 환각 완화 기법의 실용적 적용 가능성에 대한 구체적 사례는 부족하다.

실용적 활용

의료, 법률, 자동화 시스템 등 정확성과 신뢰성이 필수적인 분야에서 FMs의 활용을 안전하게 확장할 수 있도록 도와준다. 또한, 모델 개발자들이 환각을 줄이는 방향으로 모델을 설계하고, 사용자에게 신뢰도를 제공할 수 있도록 지원한다.