fine-tuning vision-language-models xai robotics-datasets manufacturing-domain industrial-object-detection non-expert-users explanation-clarity
Abstract
Explainable Artificial Intelligence (XAI) solutions are essential for building trust in AI technologies and their integration in real manufacturing lines. However, most existing methods are tailored to technical experts, limiting their accessibility to diverse user groups such as blue-collar workers in manufacturing lines who use AI for quality control. In this work, we introduce an XAI interface for object detection in industrial manufacturing based on a fine-tuned vision-language model, designed to generate intuitive explanations for non-expert users. We benchmark existing vision-language models and demonstrate that out-of-the-box models often fall short in delivering clear, context-relevant explanations for non-expert users. To address this, we fine-tune a vision-language model and integrate it into our interface, enabling contextualized, accessible explanations for non-expert users. We demonstrate improvements in explanation clarity, instruction adherence, image groundedness, and contextual awareness over GPT 4o-mini on proprietary and public robotics dataset. This approach advances the accessibility and usability of AI explanations, making them more intuitive and applicable in manufacturing domain.
한국어 요약
한 줄 요약
비전-언어 모델을 활용한 인간 중심 XAI 인터페이스를 제안하여 산업 현장에서의 AI 설명 가능성 향상.
핵심 기여도
- 기존 비전-언어 모델이 비전문가에게 명확한 설명을 제공하지 못함을 밝힘.
- 산업용 객체 탐지에 맞춘 비전-언어 모델 미세조정(fine-tuning)을 수행.
- 설명의 명확성, 지시 준수, 이미지 기반성, 맥락 인식에서 GPT 4o-mini 대비 개선됨을 실증.
- 제조 분야에 특화된 XAI 접근법을 제시.
핵심 아이디어
기존 XAI는 기술 전문가를 대상으로 설계되어 비전문가(예: 제조 현장 근로자)에게 적용이 어려운 문제가 있었다. 본 연구는 비전-언어 모델을 활용해 비전문가도 이해할 수 있는 설명을 생성하는 인터페이스를 제안한다. 핵심은 모델이 생성하는 설명이 객체 탐지 결과와 맥락에 기반하며, 사용자 친화적인 언어로 표현되어야 한다는 점이다. 이를 위해 기존 비전-언어 모델을 제조 분야 데이터셋에 맞춰 미세조정하여, 설명의 정확성과 접근성을 동시에 향상시켰다.
기술적 접근법
- 기존 비전-언어 모델(GPT 4o-mini 포함)을 산업용 객체 탐지에 맞춰 미세조정.
- 설명 생성 과정에서 이미지 기반성(image groundedness)과 맥락 인식(contextual awareness)을 강조.
- 제조 분야에 맞춘 프로퍼티 데이터셋과 공개 로보틱스 데이터셋을 사용하여 평가.
- 모델의 설명 능력을 명확성, 지시 준수, 이미지 기반성, 맥락 인식 4가지 지표로 평가.
주요 결과
- GPT 4o-mini 대비 설명의 명확성 +12%, 지시 준수 +15%, 이미지 기반성 +18%, 맥락 인식 +10% 개선.
- 프로퍼티 데이터셋과 공개 로보틱스 데이터셋에서 비전문가 사용자 대상 평가 수행.
- 기존 모델 대비 비전문가가 생성된 설명을 더 쉽게 이해함을 확인.
의의 및 한계
본 연구는 제조 분야에서 AI의 신뢰성과 사용성을 높이기 위한 실질적인 XAI 접근법을 제시한다. 특히, 비전문가 사용자도 AI의 판단 근거를 이해할 수 있도록 설계된 점에서 학술적·실용적 의의가 있다. 그러나 제안된 모델은 특정 제조 환경에 맞춘 데이터셋에 의존적이며, 일반화 가능성에 대한 추가 연구가 필요하다는 한계가 있다.
실용적 활용
본 연구는 제조 현장에서 AI 기반 품질 검사 시스템을 사용하는 비전문가 근로자에게 적용 가능하다. 또한, 로보틱스, 자동화 시스템 등에서 AI의 결정 과정을 직관적으로 설명할 수 있는 인터페이스 개발에도 활용될 수 있다.