From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection
Mike Szklarzewski, CJ George, Gavin Smithson, Christopher Stokes, Dakota Fulp, William M. Jones, Benjamin Wynn, Alexander Ur, Agit Yesiloz, Clint Kallenbach, Mark Swartz, Nathan DeBardeleben, Sharmistha Chakrabarti
arXiv:2608.07770 · 2026-08-11 공개 · arXiv · PDF
anomaly-detection unsupervised-learning human-in-the-loop defect-detection benchmark-performance validation-engine bowtie-dataset manufacturing
Abstract
Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear. In this work, we evaluate 19 unsupervised anomaly detection models on the BowTie dataset, a challenging manufacturing dataset with reflective surfaces, subtle defects, and profile-specific variation. In contrast to benchmark results, we observe that model performance is less stable than typically reported on standard benchmarks such as MVTec AD, highly sensitive to preprocessing, and inconsistent across conditions, with no single approach emerging as uniformly robust; a consensus audit further indicates that nominal-data quality affects deployment.
Motivated by these findings, we developed and initially deployed a unified human-in-the-loop framework for manufactured-part inspection that combines image annotation, AI-assisted defect detection, and an integrated validation engine, replacing a prior manual visual inspection and documentation workflow. The system supports heatmap-guided defect review, SAM-refined candidate regions for inspector acceptance, rejection, or boundary adjustment, mask evaluation where annotations exist, and review history for inspector consistency and onboarding. Together, the results highlight the gap between benchmark performance and deployment reality, and provide a practical framework for addressing it.
한국어 요약
한 줄 요약
19개의 비지도 이상 탐지 모델이 산업 현장 데이터셋에서 기대 이하의 성능을 보여, 인간-인-더-루프(HITL) 프레임워크를 제안한다.
핵심 기여도
- 19개 이상 탐지 모델을 BowTie 데이터셋에서 평가, MVTec AD 대비 불안정한 성능을 확인.
- 모델 성능이 전처리에 매우 민감하며, 조건 간 일관성 부족.
- SAM(Segment Anything Model) 기반 후보 영역 정제와 heatmap 기반 검토를 결합한 HITL 프레임워크를 제안.
- 기존 수작업 점검을 대체하는 통합 검증 엔진을 도입.
핵심 아이디어
기존 이상 탐지 모델은 MVTec AD와 같은 학술적 벤치마크에서 높은 성능을 보이지만, 실제 산업 환경인 BowTie 데이터셋에서는 예상보다 불안정하다. 이는 반사 표면, 미세 결함, 프로필별 변동성 등 실제 조건이 모델의 일반화 능력을 저하시키기 때문이다. 따라서 연구자는 HITL 프레임워크를 통해 AI와 인간 검사자의 협업을 구조화함으로써, 모델의 한계를 보완하고 검증 과정을 체계화하려는 접근을 제안한다. 특히 SAM 모델을 활용해 후보 영역을 정제하고, heatmap을 통해 결함 위치를 가이드하는 방식이 핵심이다.
기술적 접근법
- 평가 대상 모델: 19개의 비지도 이상 탐지 모델.
- 데이터셋: BowTie (반사 표면, 미세 결함, 프로필별 변동성 포함).
- 프레임워크 구성:
- Image annotation
- AI-assisted defect detection
- Integrated validation engine
- SAM(Segment Anything Model)을 사용한 후보 영역 정제.
- Heatmap 기반 결함 검토 및 검사자 피드백 반영.
- Mask evaluation 및 검토 이력 추적을 통한 일관성 관리.
주요 결과
- BowTie 데이터셋에서 모델 성능은 MVTec AD 대비 불안정하며, 전처리에 민감.
- 특정 모델이 일관된 robustness를 보이지 않음.
- HITL 프레임워크는 수작업 점검 대비 효율성과 정확도를 동시에 개선.
의의 및 한계
이 연구는 이상 탐지 모델의 학술적 성능과 실제 산업 적용 사이의 괴리를 명확히 드러내며, 이를 해소하기 위한 구조적 접근을 제시한다. 특히 SAM과 heatmap을 활용한 결함 검토는 실용적 가치가 높다. 다만, 프레임워크의 확장성이나 다양한 산업 환경에서의 적용 가능성은 추가 연구가 필요하다. 또한, BowTie 데이터셋 외 다른 데이터셋에서의 성능 검증도 명시되지 않았다.
실용적 활용
제조업 분야에서 결함 검출 프로세스 자동화 및 검증 효율화에 활용 가능하다. 특히, AI와 인간 검사자의 협업이 필요한 고정밀 검사 환경에서 유용하며, 검사자 교육 및 일관성 관리에도 기여할 수 있다.