DriftAD: Visually-Guided Text Drift for Few-Shot Industrial Anomaly Detection

Wenyang Liu, Tianyi Liu, Dongshuo Zhang, Kejun Wu, Adams Wai-Kin Kong

arXiv:2608.23723 · 2026-08-26 공개 · arXiv · PDF

vision-language-models clip mvtec-ad vis-a spatial-gating industrial-defects drift-separation-loss few-shot-anomaly-detection

Abstract

Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual features with text descriptions of normal and abnormal states. However, existing methods typically rely on static text prompts that are applied uniformly across the entire feature hierarchy and spatial dimensions. This rigid global-to-local matching fails to capture the highly localized and scale-dependent physical variations of industrial defects. To address this, we propose DriftAD, a FSAD framework built on three key modules. First, an Anomaly Signal Amplification (ASA) module enhances subtle defect signals through spatial and frequency branches before text-visual matching. Second, Visually-Guided Text Drift (VGTD) dynamically transforms frozen CLIP text embeddings, steering them into layer?wise, spatially-adaptive anomaly descriptors conditioned on local visual context at each encoder depth. Third, Drift-Guided Spatial Gating (DGSG) uses the drifted abnormal descriptor as a spatial probe to selectively enhance anomaly-relevant visual features. Addi?tionally, a drift separation loss prevents representational collapse of the drifted descriptors, and a gate supervision loss enforces spatially discriminative gating in DGSG. Extensive experiments on MVTec?AD and VisA demonstrate state-of-the-art performance across all 1-, 2-, and 4-shot settings on both image-level and pixel-level metrics. Code is available at https://github.com/wenyang001/DriftAD.

한국어 요약

한 줄 요약

DriftAD는 CLIP 기반의 시각-언어 모델을 활용한 소량 샘플 산업 이상 탐지에서 정확도를 향상시키는 시각적으로 유도된 텍스트 드리프트 기법을 제안한다.

핵심 기여도

핵심 아이디어

기존 CLIP 기반 이상 탐지 방법은 정상/이상 상태를 설명하는 텍스트 임베딩이 고정되어 있으며, 입력 이미지의 공간적 맥락이나 레이어별 특성을 반영하지 못한다는 한계가 있다. DriftAD는 이 문제를 해결하기 위해 텍스트 임베딩을 입력 이미지의 시각적 맥락에 따라 레이어별, 공간적으로 조정하는 Visually-Guided Text Drift (VGTD)를 제안한다. 이는 CLIP 텍스트 임베딩이 고정된 전역 토큰이 아닌, 입력 이미지의 각 위치와 인코더 깊이에 따라 동적으로 변형되도록 한다. 이 드리프트된 텍스트 기술자는 Drift-Guided Spatial Gating (DGSG)에서 시각적 특성을 선택적으로 강화하는 공간적 프로브 역할을 하며, 정확한 이상 지역 탐지를 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

DriftAD는 CLIP 기반 이상 탐지에서 텍스트 기술자의 공간적 및 레이어별 적응성을 도입함으로써, 기존 방법의 전역적 고정 표현의 한계를 극복한다. 특히, ASA와 VGTD를 통해 미세한 이상 신호를 강화하고, DGSG를 통해 공간적 구별력을 높이는 점에서 학술적·실용적 가치가 있다. 그러나 드리프트 기술자의 표현력이 과도하게 변형될 수 있는 문제를 방지하기 위해 드리프트 분리 손실을 도입했으며, 이는 모델의 안정성 확보에 기여한다. 한계로는 드리프트 기술자의 복잡성 증가로 인한 계산 비용 증가가 있을 수 있다.

실용적 활용

DriftAD는 제조업, 품질 검사, 자동화된 시각 검사 시스템 등에서 소량 샘플로도 정확한 이상 탐지를 요구하는 산업 분야에 적용 가능하다. 특히, 다양한 제품 카테고리에 대한 모델 재훈련 없이도 일반화 가능한 이상 탐지가 필요한 상황에서 유용하다.