in-context-learning few-shot-learning clip residual-learning industrial-defects visual-language-models generalist-anomaly-detection medical-anomalies
Abstract
This paper explores the problem of Generalist Anomaly Detection (GAD), aiming to train one single detection model that can generalize to detect anomalies in diverse datasets from different application domains without any further training on the target data. Some recent studies have showed that large pre-trained Visual-Language Models (VLMs) like CLIP have strong generalization capabilities on detecting industrial defects from various datasets, but their methods rely heavily on handcrafted text prompts about defects, making them difficult to generalize to anomalies in other applications, e.g., medical image anomalies or semantic anomalies in natural images. In this work, we propose to train a GAD model with few-shot normal images as sample prompts for AD on diverse datasets on the fly. To this end, we introduce a novel approach that learns an in-context residual learning model for GAD, termed InCTRL. It is trained on an auxiliary dataset to discriminate anomalies from normal samples based on a holistic evaluation of the residuals between query images and few-shot normal sample prompts. Regardless of the datasets, per definition of anomaly, larger residuals are expected for anomalies than normal samples, thereby enabling InCTRL to generalize across different domains without further training. Comprehensive experiments on nine AD datasets are performed to establish a GAD benchmark that encapsulate the detection of industrial defect anomalies, medical anomalies, and semantic anomalies in both one-vs-all and multi-class setting, on which InCTRL is the best performer and significantly outperforms state-of-the-art competing methods. Code is available at https://github.com/mala-lab/InCTRL.
한국어 요약
한 줄 요약
InCTRL은 few-shot 정상 이미지를 기반으로 잔차를 학습하여 다양한 도메인의 이상 탐지를 일반화하는 새로운 GAD 모델이다.
핵심 기여도
- GAD(Generic Anomaly Detection)라는 새로운 평가 태스크를 제안하여, 타겟 데이터셋에 대한 학습 없이 다양한 도메인의 이상을 탐지하는 일반화 능력을 평가.
- InCTRL이라는 in-context residual learning 기반의 GAD 모델을 제안, CLIP 기반 잔차 분석을 통해 정상/이상 구분.
- 9개 이상 탐지 데이터셋에서 one-vs-all 및 multi-class 설정에서 기존 SOTA 모델 대비 우수한 성능을 달성.
- few-shot 정상 샘플을 inference 시에만 사용하여 모델 학습 과정에서 과적합을 방지.
핵심 아이디어
기존 CLIP과 같은 VLM(Vision-Language Model)은 정교한 텍스트 프롬프트를 필요로 하여 의료 이미지나 의미적 이상 탐지 등 다른 도메인으로 확장이 어려웠다. InCTRL은 few-shot 정상 이미지를 in-context sample prompts로 활용하여, 잔차(residual)를 기반으로 정상/이상을 구분하는 새로운 접근법을 제시한다. 이는 CLIP의 잔차를 이미지 및 패치 수준에서 분석하여, 다양한 도메인에서 일반화 가능한 이상 탐지를 가능하게 한다. InCTRL은 정상 샘플과의 잔차가 이상 샘플일수록 더 커진다는 점을 활용하여, 학습 과정에서 타겟 데이터셋에 의존하지 않고도 일반화된 이상 탐지를 수행한다.
기술적 접근법
- **InCTRL 모델**: CLIP 기반 잔차 학습 모델로, 정상 샘플과의 잔차를 기반으로 이상 탐지를 수행.
- **In-context residual learning**: 이미지 및 패치 수준에서 잔차를 분석하여, 정상/이상 구분.
- **Few-shot sample prompts**: 정상 샘플을 inference 시에만 사용, 학습 과정에서 사용하지 않음.
- **Auxiliary dataset**: InCTRL은 보조 데이터셋에서 정상/이상 구분을 학습하며, 타겟 데이터셋에 대한 학습 없이도 일반화.
- **Text prompt 활용**: 텍스트-이미지 정렬된 의미 공간에서 정상/이상 텍스트 프롬프트를 결합하여 성능 강화.
주요 결과
- InCTRL은 9개 이상 탐지 데이터셋에서 one-vs-all 및 multi-class 설정에서 기존 SOTA 모델 대비 **유의미한 성능 개선**을 보임.
- **Industrial defect, medical, semantic 이상 탐지** 모두에서 우수한 성능을 달성.
- few-shot 설정에서도 모델 성능이 안정적으로 유지됨.
- 정확도 수치는 명시되지 않지만, **"significantly outperforms state-of-the-art competing methods"**로 명시됨.
의의 및 한계
InCTRL은 기존 이상 탐지 모델이 타겟 데이터셋에 의존하는 문제를 해결하며, few-shot 정상 샘플만으로도 일반화된 이상 탐지를 가능하게 한다는 점에서 학술적·실용적 의의가 크다. 특히, 텍스트 프롬프트에 의존하지 않는 방식은 의료 이미지나 의미적 이상 탐지 등 다양한 도메인으로 확장 가능성을 열어준다. 그러나, few-shot 샘플의 질이나 수에 따라 성능이 변동할 수 있으며, 특정 도메인에서의 잔차 분석 한계가 있을 수 있다. 또한, 정확한 성능 수치(예: AUC, F1-score 등)는 명시되지 않아 비교의 한계가 있다.
실용적 활용
InCTRL은 데이터 수집이 어려운 의료 분야나 데이터 프라이버시가 중요한 산업 검사 분야에서 유용하게 활용될 수 있다. 또한, 신속한 이상 탐지가 필요한 실시간 시스템에서 few-shot 샘플만으로 모델을 적용할 수 있어, 신속한 배포와 유지보수를 가능하게 한다.