Asymptotic Risk Calibration for Selective Question Answering
This work addresses the challenge that large language models often produce fluent yet incorrect answers in question answering, while existing uncertainty scoring methods struggle to reliably distinguish correct from incorrect responses and fail to statistically control error rates with fixed thresholds. To overcome these limitations, the authors propose A-CRC-QA, a novel framework that introduces asymptotic risk calibration into selective question answering for the first time. Inspired by conformal risk control, their approach employs a monotonic empirical risk calibration procedure that reformulates error control as a linear expectation constraint. The method is training-free, model-agnostic, compatible with diverse uncertainty estimators, and applicable to both open-ended and closed-form QA tasks. Experiments on CoQA and MedMCQA demonstrate that A-CRC-QA achieves a significantly better trade-off between answer retention rate and reliability compared to uncalibrated baselines and confidence-bound approaches.