신뢰도는 바뀌고, 예측은 바뀌지 않는다: 사후 보정을 위한 예측 보존 수리
Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
September 2, 2026
저자: Daehwan Kim, Haejun Chung, Ikbeom Jang
cs.AI
초록
사후 보정은 보고된 신뢰도를 교정하지만, 다중 클래스 보정기는 연관된 top-1 예측까지 변경할 수 있다. 정확도는 이러한 변경이 정답 여부에 미치는 순효과만 포착할 뿐 예측이 얼마나 자주 변경되는지는 포착하지 못하는 반면, top-1 예측 변경률(TPCR)은 이 빈도를 측정한다. 우리는 전체 보정 확률 벡터를 수리함으로써 정확한 예측 보존을 부과하는 최초의 적합 후 어댑터인 CORD(Top-1 결정 보존을 위한 보정기 출력 수리, Calibrator-Output Repair for Top-1 Decision Preservation)를 제안한다. CORD는 원래 출력과 보정 출력만으로 원래 top-1에 할당된 질량을 결정한다. 보정된 조건부 분포는 나머지 질량을 다른 클래스들에 배분하며, 그 결과 얻어진 수리된 벡터는 자체의 argmax가 원래 예측을 복원한다. 보정 분할에서 CORD는 달성 가능할 때마다 원래 예측에 대한 보정 출력의 평균 질량을 유지하도록 수리된 질량들을 조정한다. 이 어댑터는 적합된 보정기나 그 직접 출력을 변경하지 않으며, 추가적인 지도 학습 맵을 적합시키지 않고, 사용자 또는 검증으로 조정되는 하이퍼파라미터도 요구하지 않는다. CIFAR-10/100과 ImageNet-1K 전반에서 CORD는 구조적으로 TPCR이 0임을 달성하며, 모든 데이터셋에서 해당 직접 출력 대비 평균 ECE, NLL, Brier 점수를 낮춘다. 이러한 쌍별 비교에서의 개선은 분포 이동 및 다양한 보정 집합 크기에서도 유지된다. 따라서 CORD는 보존 제약을 보정기 적합 과정에서 제거하고, 원래 결정의 정확한 복원을 후속 출력 수리에 할당한다. 우리의 코드는 https://github.com/labhai/CORD에서 확인할 수 있다.
English
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.