Distribution-Free Inference for the Regression Function of Binary Classification

📅 2023-08-03
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of conventional regression function inference—namely, its reliance on distributional assumptions and asymptotic approximations. We propose the first fully distribution-free, non-asymptotic framework for constructing confidence regions for regression functions. Our method leverages resampling techniques without requiring parametric modeling assumptions and is valid for any finite sample size and user-specified confidence level. Theoretically, we establish strong consistency and PAC-style exclusion bounds. Notably, this is the first method to deliver exact, finite-sample confidence sets for conditional expectation functions in binary classification, enabling reliable construction of Bayes-optimal classifiers. Numerical experiments demonstrate superior performance: our approach achieves more robust coverage and tighter intervals—particularly in small-sample regimes and under non-regular distributions (e.g., non-Gaussian, heavy-tailed)—outperforming traditional asymptotic ellipsoidal methods.
📝 Abstract
One of the key objects of binary classification is the regression function, i.e., the conditional expectation of the class labels given the inputs. With the regression function not only a Bayes optimal classifier can be defined, but it also encodes the corresponding misclassification probabilities. The paper presents a resampling framework to construct exact, distribution-free and non-asymptotically guaranteed confidence regions for the true regression function for any user-chosen confidence level. Then, specific algorithms are suggested to demonstrate the framework. It is proved that the constructed confidence regions are strongly consistent, that is, any false model is excluded in the long run with probability one. The exclusion is quantified with probably approximately correct type bounds, as well. Finally, the algorithms are validated via numerical experiments, and the methods are compared to approximate asymptotic confidence ellipsoids.
Problem

Research questions and friction points this paper is trying to address.

Construct distribution-free confidence regions for regression function in binary classification
Ensure strong uniform consistency for empirical risk minimization in finite pseudo-dimension models
Provide exponential PAC bounds on L2 sizes of confidence regions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Resampling test for distribution-free confidence regions
Empirical risk minimization with pseudo-dimensions
Exponential bounds on L2 region sizes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
SZTAKI | ELKH | Eötvös Loránd University
A
Ambrus Tamás
SZTAKI: The Institute for Computer Science and Control, ELKH: Eötvös Loránd Research Network, 13-17 Kende utca, 1111, Budapest, Hungary; Institute of Mathematics, Eötvös Loránd University (ELTE), Budapest, Hungary
Balázs Csanád Csáji
Balázs Csanád Csáji
SZTAKI: Institute for Computer Science and Control, Budapest, Hungary
machine learningsystem identificationcontrol theoryoptimizationstatistics