🤖 AI Summary
This study addresses the challenge of modeling extreme tail regression for high-dimensional response variables by proposing a high-dimensional extreme tail generalized linear model. By integrating Bregman divergence with L1 penalization for parameter estimation and constructing a tail-localized debiased estimator, this work establishes a unified inference framework applicable to diverse tail types. Theoretically, we prove the convergence rate and asymptotic normality of the estimator, thereby enabling valid confidence interval construction even when covariate dimensionality exceeds sample size. This approach effectively facilitates statistical inference for high-dimensional extreme data and is successfully validated through an empirical analysis of automobile insurance claims.
📝 Abstract
We propose a regression model for the extreme tail of a response variable, in which covariates rescale the tail without changing its shape. A single covariate-dependent function then characterizes the entire conditional tail, in contrast to extreme quantile regression, which targets a quantile at a pre-specified level. The tail shape itself is left unrestricted: heavy-, light- and short-tailed responses are covered by the same framework. We specify the function through a link function and a linear combination of the covariates, which is in the spirit of a generalized linear model. In estimation, we match the parametric specification to the underlying tail function under a Bregman divergence, over a region localized at the largest observations. The resulting loss is convex, and an $\ell_1$-penalty allows the number of covariates to exceed the effective sample size. The tail localization makes the asymptotic theory deviate from that for classical penalized generalized linear models. Only the tail observations selected by a random threshold are used in the statistical analysis, making them dependent. We derive the convergence rate of the penalized estimator and propose a debiased estimator that is asymptotically normal, yielding confidence intervals for individual coefficients. Its asymptotic variance is determined by the covariance of the score, which under tail localization differs from the Hessian and must be estimated separately. We apply the method to automobile insurance claims data.