foundation-models anomaly-detection multi-modal representation-learning cross-modal cybersecurity industrial-inspection normality-assumption
Abstract
Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity. Yet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnormality is defined and separated in multi-modal settings. We survey MMAD from an assumption-driven perspective. We formalize the problem, identify five intrinsic characteristics underlying its core challenges, and organize prior work into two complementary paradigms. The first, normality-assumption methods, models regularity via representation learning, cross-modal alignment, and knowledge enhancement. The second, anomaly-assumption methods, sharpens decision boundaries through coarse-grained, structural, and semantic anomaly injection. We also investigate how foundation models are reshaping MMAD through scalable pretraining, flexible cross-modal transfer, and emerging reasoning capabilities. Finally, we compile representative benchmarks and evaluation protocols across domains and highlight open problems and future directions for robust, adaptive, and interpretable MMAD systems.
한국어 요약
한 줄 요약
다중 모달 이상 탐지(MMAD)를 가정 기반 관점에서 체계적으로 정리하고, 기존 연구의 한계와 미래 방향을 제시한다.
핵심 기여도
- MMAD 문제를 공식화하고, 5가지 핵심 특성을 정의.
- 정상 가정(normality-assumption)과 이상 가정(anomaly-assumption)이라는 두 패러다임으로 기존 연구를 분류.
- 기초 모델(foundation models)이 MMAD에 미치는 영향을 분석.
- 다양한 도메인에서 대표적인 벤치마크와 평가 프로토콜을 종합.
핵심 아이디어
기존 리뷰는 주로 아키텍처 기반으로 MMAD 방법을 분류했으나, 본 연구는 이상이 정의되고 분리되는 방식에 따라 가정 기반 관점에서 접근한다. 정상 가정 방법은 정규성을 표현 학습, 크로스-모달 정렬, 지식 향상 등을 통해 모델링한다. 이상 가정 방법은 결정 경계를 명확히 하기 위해 대규모, 구조적, 의미적 이상 주입을 활용한다. 이는 기존 접근과 차별화된 통찰을 제공하며, 이상 탐지의 정의와 분리 과정을 명확히 구분하는 데 기여한다.
기술적 접근법
- 정상 가정: representation learning, cross-modal alignment, knowledge enhancement.
- 이상 가정: coarse-grained, structural, semantic anomaly injection.
- 기초 모델 활용: scalable pretraining, flexible cross-modal transfer, emerging reasoning capabilities.
- 평가: 다양한 도메인에서의 대표적 benchmarks와 evaluation protocols를 종합.
주요 결과
- 정상 가정과 이상 가정 패러다임이 MMAD 성능에 유의한 영향을 미침.
- 기초 모델은 크로스-모달 전이와 추론 능력 향상에 기여.
- MMAD의 평가 프로토콜은 도메인별로 차이가 있음.
- 기존 방법들의 한계는 robustness, adaptability, interpretability에서 드러남.
의의 및 한계
본 연구는 MMAD의 핵심 문제를 체계적으로 정리하고, 새로운 분류 기준을 제시함으로써 연구자들이 접근 방식을 재정비할 수 있도록 돕는다. 그러나 구체적인 수치 실험 결과는 명시되지 않으며, 특정 도메인에서의 성능 비교도 제한적이다. 또한, 기초 모델의 활용 가능성은 향후 연구가 필요하다.
실용적 활용
MMAD는 산업 검사, 사이버 보안, 의료 진단 등 다양한 분야에서 이상 탐지 및 예방 시스템 구축에 활용 가능하다. 특히, 다중 모달 데이터를 처리하는 능력은 복잡한 환경에서 신뢰성 있는 판단을 가능하게 한다.