xgboost random-forest ensemble-learning logistic-regression distributional-shift bearing-fault-detection entropy-estimation neutrosophic-classification
Abstract
Machine learning classifiers for bearing fault detection produce scalar confidence scores that conflate confident errors with genuinely ambiguous predictions, and the conventional truth/falsity pair (F = 1 - T) is algebraically redundant by construction. We operationalize a refined neutrosophic decomposition of a Random Forest + XGBoost + Logistic Regression ensemble into four indicators -- T-hat (top-class evidence), F-hat (best-competitor evidence), predictive entropy I1-hat, and decision disagreement I2-hat -- evaluated on two bearing benchmarks (CWRU and JNU, 600-1000 rpm) under a leave-one-condition-out protocol. On CWRU, after correcting a file-to-class mapping error, the ensemble reaches 100.00 percent accuracy on three of four held-out loads (92.27 percent on the fourth), leaving too few errors for uncertainty analysis. On JNU, holding out 1000 rpm, accuracy collapses to 40.64 percent, below a majority-class baseline; Logistic Regression (57.91 percent) generalizes far better than the tree ensembles. I1-hat shows a robust association with error beyond T-hat/F-hat, while I2-hat contributes little; standalone Logistic Regression confidence outperforms the full decomposition, a boundary condition we report honestly. Two further results extend this: fusing a time-domain and a frequency-domain model of the same signal and scoring their Jensen-Shannon divergence beats that model own entropy (AURC 0.29 vs. 0.36 on the standard split; 0.54 vs. 0.73 under a harder single-condition reproduction), the only indicator moving correctly under a CWRU-versus-JNU distributional-shift contrast; and, on CWRU alone, literature-verified bearing fault frequencies, correctly demodulated via the envelope spectrum, separate most fault classes almost perfectly (99.57 percent) using three interpretable features. Code, logs, and figures are released for independent verification.
한국어 요약
한 줄 요약
베어링 고장 탐지에서 불확실성 인식을 위한 중성론적 앙상블 분류 방법을 제안하고 실험적으로 검증한다.
핵심 기여도
- Random Forest + XGBoost + Logistic Regression 앙상블을 중성론적 분해로 4개 지표(T-hat, F-hat, I1-hat, I2-hat)로 분리.
- CWRU 데이터셋에서 3개 조건에서 100.00% 정확도 달성.
- JNU 데이터셋에서 Logistic Regression이 트리 앙상블보다 17.27% 높은 정확도(57.91%) 보임.
- 시간 영역과 주파수 영역 모델의 Jensen-Shannon 발산이 모델 엔트로피보다 AURC 0.07~0.19 개선.
핵심 아이디어
기존 이진 신뢰도(T/F)는 수학적으로 중복되며, 오류와 불확실성을 구분하지 못한다. 본 연구는 중성론적 분해를 통해 예측의 불확실성을 4개 지표로 구분: T-hat(최상위 클래스 증거), F-hat(최대 경쟁자 증거), I1-hat(예측 엔트로피), I2-hat(결정 불일치). 이는 앙상블 모델의 신뢰도를 더 세분화하여 분석할 수 있도록 한다. 특히, I1-hat은 오류와의 상관성이 T-hat/F-hat보다 높았으며, I2-hat은 기여도가 낮았다. 또한, 독립적인 모델 간의 Jensen-Shannon 발산이 불확실성 측정에 효과적임을 보여준다.
기술적 접근법
- 모델: Random Forest, XGBoost, Logistic Regression 앙상블.
- 데이터: CWRU, JNU 베어링 데이터셋 (600-1000 rpm).
- 평가: leave-one-condition-out 프로토콜.
- 지표: T-hat, F-hat, I1-hat, I2-hat.
- 추가 실험: 시간 영역과 주파수 영역 모델의 Jensen-Shannon 발산 계산.
주요 결과
- CWRU: 3개 조건에서 100.00% 정확도, 4번째 조건에서 92.27%.
- JNU: 1000 rpm 제외 시 40.64% 정확도, Logistic Regression은 57.91% (기본값 대비 +17.27%).
- Jensen-Shannon 발산: 표준 분할에서 AURC 0.29 (기존 0.36), 단일 조건 재현 시 0.54 (기존 0.73).
- CWRU에서 고장 주파수 복조(엔벨로프 스펙트럼)로 99.57% 정확도 달성.
의의 및 한계
- 중성론적 분해는 불확실성 분석을 구조화하며, 앙상블 모델의 신뢰도를 더 정확히 평가할 수 있다.
- 그러나 I2-hat의 기여도가 낮고, Logistic Regression의 단독 신뢰도가 앙상블 분해보다 우수한 경우가 있어 한계가 있음.
- 데이터셋 간 분포 차이(CWRU vs. JNU)에서 모델의 일반화 능력이 제한적임을 보여준다.
실용적 활용
베어링 고장 탐지 시스템에서 불확실성 인식이 필요한 산업 현장(예: 제조, 에너지)에 적용 가능하며, 모델 신뢰도를 정량적으로 평가하여 유지보수 결정을 지원할 수 있다.