ReliableNet: A Chance-Constrained Approach to Trustworthy Classification in Deep Learning
This work addresses the critical reliability challenge in deep classification known as “high-confidence yet wrong” (JCW) predictions, which undermine abstention and human-in-the-loop mechanisms. It proposes the first training-stage formulation of trustworthy classification as a chance-constrained empirical risk minimization problem, directly bounding the JCW probability by a user-specified risk budget α via a conservative smooth inner approximation. This approach yields a theoretically sound smooth surrogate that enables rigorously certified control over JCW risk. Empirical evaluations demonstrate that the proposed ReliableNet consistently adheres to the prescribed JCW budget across four tabular and two image datasets, achieves the lowest empirical JCW under distribution shift, and excels in accuracy, coverage, calibration, and selective prediction performance.