Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

Filippo Cenacchi, Longbing Cao, Runze Yang

arXiv:2607.18279 · 2026-07-22 공개 · arXiv · PDF

time-series-classification spectral-evidence reliability-estimation post-hoc-calibration ucr-uea-datasets validation-gating band-energy selective-reliability

Abstract

Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal. We address three time-series reliability gaps: identical confidence values can hide different temporal support, average calibration can miss false high-confidence errors, and output-space recalibration offers limited input-linked auditability. We introduce a validation-gated fixed-label reliability policy that keeps the backbone prediction unchanged while estimating whether it should be trusted. The method combines output-side cues with whole-sample spectral descriptors, including band energy, entropy, peak dominance, period support, and phase stability, to form a scalar reliability estimate and diagnostic band-level evidence. A validation gate enables spectral conditioning only when correctness ranking improves without breaching [email protected] or AURC tolerances; otherwise it reverts to the safer output-space baseline. Across eight heterogeneous UCR/UEA datasets, eight time-series backbone families, and standard recalibrators, the unconstrained method improves fixed-label selective-reliability metrics on the matched evaluation subset, raising Corr-AURC from 0.693 to 0.779. The validation-gated policy further improves Corr-AURC to 0.786 and reduces [email protected] to 0.094. These results suggest that reliability estimation for time-series classifiers benefits from bundling output confidence with spectral evidence, while validation gating prevents unsupported spectral conditioning.

한국어 요약

한 줄 요약

SEB-Cal은 시간 시계열 분류에서 신뢰도와 스펙트럼 증거를 결합하여 선택적 신뢰성 추정을 개선하는 검증 게이트 기반 정책을 제안한다.

핵심 기여도

핵심 아이디어

기존의 시간 시계열 분류에서 사후 신뢰도 조정은 주로 출력 공간에서 이루어지며, 이는 신뢰도가 실제 신뢰성을 반영하지 못하는 경우가 많다. 예를 들어, 동일한 0.92의 신뢰도는 전역적으로 조직된 신호에서 나올 수도 있고, 노이즈나 국지적 특징에서 나올 수도 있다. 이 연구는 **전체 샘플의 스펙트럼 특성**(에너지, 엔트로피, 피크 지배, 위상 안정성)을 고려하여 신뢰도를 보완적으로 평가하는 새로운 접근법을 제안한다. 이는 단순히 출력 공간의 확률을 재정의하는 방식과는 달리, 입력 신호 자체의 전역적 조직성을 평가하는 **스펙트럼 증거 결합**(Spectral Evidence Bundling)을 기반으로 한다. 특히, **SEB-Cal**은 예측 라벨을 변경하지 않으면서 신뢰도를 재평가하는 정책으로, **validation gate**를 통해 안전 기준을 만족하는 경우에만 스펙트럼 조건을 활성화한다.

기술적 접근법

주요 결과

의의 및 한계

SEB-Cal은 시간 시계열 분류에서 신뢰도가 단순히 출력 공간의 확률만 기반하지 않고, 입력 신호의 전역적 조직성까지 고려하는 새로운 접근법을 제시한다. 특히, **validation gate**를 통해 안전 기준을 만족하는 경우에만 스펙트럼 조건을 활성화함으로써, 과도한 신뢰도 상승을 방지하는 데 기여한다. 이는 시간 시계열 분류에서 신뢰성 추정을 **evidence-constrained reliability selection** 문제로 재정의하는 데 기여한다. 그러나 모든 데이터셋과 백본에서 유의한 개선이 관찰되지 않았으며, 일부 설정에서는 **validation gate**에 의해 스펙트럼 조건이 거부되는 경우가 발생했다. 이는 스펙트럼 증거가 모든 상황에서 유용하지 않다는 점을 시사한다.

실용적 활용

SEB-Cal은 의료, 금융, 산업 모니터링 등 시간 시계열 신뢰도가 중요한 분야에서 적용 가능하다. 특히, **신뢰도가 높은 예측이 실제 신뢰성을 반영하는지**를 평가해야 하는 상황에서 유용하며, **[email protected]**와 같은 안전 기준을 준수하면서 신뢰도를 개선할 수 있다. 이는 시스템이 신뢰도에 따라 **신뢰, 중단, 검토** 결정을 내리는 데 기여할 수 있다.