🤖 AI Summary
This work addresses the limitations of conventional regression function inference—namely, its reliance on distributional assumptions and asymptotic approximations. We propose the first fully distribution-free, non-asymptotic framework for constructing confidence regions for regression functions. Our method leverages resampling techniques without requiring parametric modeling assumptions and is valid for any finite sample size and user-specified confidence level. Theoretically, we establish strong consistency and PAC-style exclusion bounds. Notably, this is the first method to deliver exact, finite-sample confidence sets for conditional expectation functions in binary classification, enabling reliable construction of Bayes-optimal classifiers. Numerical experiments demonstrate superior performance: our approach achieves more robust coverage and tighter intervals—particularly in small-sample regimes and under non-regular distributions (e.g., non-Gaussian, heavy-tailed)—outperforming traditional asymptotic ellipsoidal methods.
📝 Abstract
One of the key objects of binary classification is the regression function, i.e., the conditional expectation of the class labels given the inputs. With the regression function not only a Bayes optimal classifier can be defined, but it also encodes the corresponding misclassification probabilities. The paper presents a resampling framework to construct exact, distribution-free and non-asymptotically guaranteed confidence regions for the true regression function for any user-chosen confidence level. Then, specific algorithms are suggested to demonstrate the framework. It is proved that the constructed confidence regions are strongly consistent, that is, any false model is excluded in the long run with probability one. The exclusion is quantified with probably approximately correct type bounds, as well. Finally, the algorithms are validated via numerical experiments, and the methods are compared to approximate asymptotic confidence ellipsoids.