Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

Daehwan Kim, Haejun Chung, Ikbeom Jang

arXiv:2609.01072 · 2026-09-06 공개 · arXiv · PDF

cifar-10 confidence-calibration imagenet-1k post-hoc-calibration prediction-preservation top-1-prediction ece-reduction probability-repair

Abstract

Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.

한국어 요약

한 줄 요약

CORD는 예측을 유지하면서 신뢰도만 보정하는 후처리 보정 기법으로, TPCR 0%와 ECE, NLL, Brier 점수 개선을 달성한다.

핵심 기여도

핵심 아이디어

기존의 다중 클래스 보정기(multiclass calibrator)는 신뢰도를 수정하면서 동시에 최상위 예측(top-1 prediction)을 바꾸는 경우가 있다. 이는 정확도(Accuracy)가 예측 변경의 빈도를 반영하지 않기 때문에, 예측 변경의 전체 발생률(TPCR)을 따로 측정해야 한다는 문제를 유발한다. CORD는 이 문제를 해결하기 위해, 기존 보정기의 출력 벡터를 수정하여 최상위 예측을 유지하면서 신뢰도만 보정하는 새로운 후처리 접근법을 제안한다. CORD는 원래 예측 클래스에 할당된 확률 질량을 복원하고, 나머지 클래스에 대한 보정된 조건부 분포를 유지함으로써, 최종 argmax가 원래 예측과 일치하도록 보장한다. 이는 기존 보정기의 파라미터나 출력을 변경하지 않으면서도, 예측 보존을 강제하는 방식이다.

기술적 접근법

주요 결과

의의 및 한계

CORD는 기존 보정기의 예측 변경 문제를 후처리 단계에서 해결함으로써, 보정기 학습 시 예측 보존 제약을 제거할 수 있는 새로운 설계 공간을 열었다. 이는 보정기의 표현력(expressiveness)을 유지하면서도, 예측 변경 여부를 사용자 결정에 따라 선택할 수 있게 한다는 점에서 실용적 가치가 있다. 다만, CORD는 단순히 확률 질량을 재할당하는 방식으로, 복잡한 보정 목적이나 다중 예측 보존(multi-label) 상황에서는 한계가 있을 수 있다. 또한, CORD는 기존 보정기의 출력을 수정하는 방식이기 때문에, 보정기 자체의 성능이 낮다면 CORD도 제한된 개선 효과를 보일 수 있다.

실용적 활용

CORD는 의료, 자율주행, 금융 등 예측 변경이 허용되지 않는 분야에서 유용하게 사용될 수 있다. 예측 결과가 반드시 유지되어야 하는 상황에서, 신뢰도만 보정하여 모델 신뢰성을 높이는 데 적합하다. 또한, 보정기 학습 시 예측 보존 제약 없이 자유로운 보정이 가능하다는 점에서, 모델 개발 및 배포 과정에서 유연성을 제공한다.