Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees

📅 2026-07-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the unreliability of large language models (LLMs) in question answering and their lack of statistically guaranteed uncertainty quantification. To this end, the authors propose the CIC framework, which provides the first finite-sample, provably aligned selective answering mechanism for LLMs. By calibrating arbitrary uncertainty scores using Hoeffding or Clopper–Pearson confidence intervals, CIC constructs response policies that strictly control the error rate under a user-specified risk level. Empirical evaluation across seven prominent LLMs and multiple uncertainty estimators demonstrates that CIC consistently maintains high answer rates while rigorously adhering to the prescribed risk constraints.
📝 Abstract
Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they may generate hallucinated or misaligned responses without reliable confidence estimates. Uncertainty quantification (UQ) offers a natural basis for selective answering, where a system answers only when its prediction is deemed reliable and abstains otherwise. However, existing uncertainty scores for LLMs are often heuristic: a threshold chosen on such scores does not, by itself, provide statistical guarantees on the error rate among accepted answers. We propose CIC, a confidence-interval-based calibration framework that converts arbitrary uncertainty scores into risk-controlled selective answering rules. Given a held-out calibration set, CIC evaluates each generated response using an application-specific alignment criterion and associates it with an uncertainty score and a binary error label. For each candidate uncertainty threshold, CIC estimates the acceptance-conditioned error rate and constructs a high-probability upper confidence bound using either Hoeffding-style or Clopper-Pearson confidence intervals. It then selects the largest threshold whose upper bound is below a user-specified risk level $α$, thereby maximizing the answering rate subject to a finite-sample reliability constraint. Under exchangeability, CIC guarantees with probability at least $1-δ$ that the selected threshold, if non-null, controls the error rate among accepted answers at level $α$. We evaluate CIC on both closed-ended and open-ended QA benchmarks across seven LLMs and multiple uncertainty estimators. Experimental results show that CIC consistently achieves valid risk control while retaining strong answering efficiency, providing a practical and statistically grounded mechanism for deploying LLMs in reliability-sensitive QA workflows.
Problem

Research questions and friction points this paper is trying to address.

uncertainty quantification
abstention
risk control
large language models
selective answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty Quantification
Selective Answering
Confidence Interval Calibration
Risk-Controlled Inference
Large Language Models
💼 Related Jobs
No related jobs found.
S
Sijin Dong
Ibaraki University, Japan
H
Hiroyuki Shinnou
Ibaraki University, Japan