Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis

Yuntong Chen, Jianyu Liu, Guobin Zhao, Ziang Wang, Chao Chen, Ju Huang, Xitian Tian, Lijiang Huang

arXiv:2607.22797 · 2026-07-28 공개 · arXiv · PDF

qlora temporal-localization frequency-estimation llm-reporting bearing-fault-diagnosis diagnostic-evidence-network validation-signal label-free-validation

Abstract

Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against physical reality before it is acted upon. Current intelligent fault diagnosers fail this standard in two ways. Their standard output, a class label with a softmax confidence score, is an internal statistic of the classifier, offering nothing checkable against independent physical knowledge; and the growing use of generative language models in maintenance reporting adds a second risk: hallucinated content entering reports on which decisions rest. Taking bearing fault diagnosis as the testbed, this work addresses both problems from the output side. The proposed Diagnostic Evidence Network (DENet) is an encoder-agnostic multi-task framework extending the output to a structured evidence record: the classification, a predicted characteristic frequency comparable against the theoretical value determined by bearing geometry and shaft speed, and a temporal localization of transient impulses inspectable on the raw waveform. Across four encoders and three public datasets, this evidence incurs no statistically significant accuracy cost, with a frequency error of about 6 Hz on 1,024-point segments where spectral estimation is structurally inapplicable. Centrally, the deviation between predicted and theoretical frequency constitutes a label-free, inference-time validation signal: it detects misclassifications with AUROC values of 0.970 and 0.871, and remains discriminative in the high-confidence regime where confidence-derived detectors are blind. Finally, a QLoRA-adapted language model is constrained to translate, but never generate, diagnostic content, reducing unsupported-claim rates from 10-12% to 2% and eliminating fabricated quantities.

한국어 요약

한 줄 요약

안전-critical 시스템에서 신뢰성 있는 AI 진단을 위해 물리적 검증 가능한 증거와 LLM 기반 보고를 결합한 새로운 접근법을 제시한다.

핵심 기여도

핵심 아이디어

기존 진단 시스템은 분류 결과만 제공하며, 이는 물리적 현실과 비교할 수 없는 내부 통계량이다. 본 연구는 DENet이라는 다중 태스크 프레임워크를 통해 분류 외에도 예측 주파수와 임펄스 위치를 추가하여, 이론적 물리 모델과 비교 가능한 증거를 생성한다. 특히, 예측 주파수와 이론적 주파수 간 편차를 실시간 검증 신호로 활용함으로써, 고신뢰도 영역에서도 오분류를 탐지할 수 있다. 또한, LLM을 진단 내용 번역에만 활용함으로써 hallucination을 최소화하는 접근법이 핵심이다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 AI 진단의 신뢰도를 높이기 위해 물리적 검증 가능한 출력 구조를 제안하며, 특히 고신뢰도 영역에서도 오분류를 탐지할 수 있는 새로운 검증 신호를 도입한 점에서 학술적 의의가 있다. 또한, LLM의 hallucination 문제를 해결함으로써 유지보수 보고서의 신뢰도를 높였다. 그러나 특정 인코더나 데이터셋에 대한 일반화 가능성은 명시되지 않았으며, 실제 산업 환경에서의 적용 가능성은 추가 연구가 필요하다.

실용적 활용

이 연구는 항공기, 자동차, 제조 설비 등 안전-critical 시스템의 유지보수 및 고장 진단에 적용 가능하다. 특히, AI 기반 진단 결과의 신뢰도를 높이고, 유지보수 보고서의 정확성을 보장하는 데 유용할 수 있다.