🤖 AI Summary
This paper addresses the challenge of accurately identifying racial disparities in loan approval decisions under regulatory fairness requirements—despite the absence of explicit race information. We propose a frequentist identifiable estimation framework that, under weak exogeneity assumptions, constructs the first consistent estimator for the Black–White approval rate differential. Our approach innovatively integrates surnames as proxy variables with income-stratified demographic priors, thereby mitigating systematic biases inherent in conventional heuristic methods. By jointly modeling via ordinary least squares (OLS) and maximum likelihood estimation (MLE), we combine surname-based racial inference with geodemographic probability models to yield robust individual-level racial probability proxies. Empirical evaluation on Los Angeles Home Mortgage Disclosure Act (HMDA) data demonstrates that our method reduces the root-mean-square error (RMSE) of the Black–White adverse impact ratio by 79.7%, substantially outperforming existing weighted estimation approaches.
📝 Abstract
Estimating racial disparities in loan-approval probabilities when race is unobserved is routinely required for fair lending compliance. In such cases, race probabilities-typically from Bayesian Improved Surname Geocoding (BISG)-stand in for true race. Prior work shows that common heuristic approaches, including the Threshold and Weighting estimators, are inconsistent under valid identification assumptions, compromising internal validity. A recent Bayesian approach demonstrates consistency under assumptions reasonable in many fair lending contexts. This approach hinges on the insight that identification requires the race predictors to be exogenous with respect to loan approval, essentially an instrumental-variables design. We present a frequentist counterpart to this solution via Ordinary Least Squares (OLS) and Maximum Likelihood Estimation (MLE) under a similar exogeneity assumption. To satisfy these assumptions in practice, we introduce (i) a surname-only proxy analogous to BISG and (ii) an income-stratified prior for race probabilities. Monte Carlo simulations and an application to 2023 Los Angeles HMDA data confirm superior performance: this method reduces RMSE in the LA Black/White adverse-impact ratio by 79.7% (from 10.639pp to 2.158pp) compared to a Weighting estimator with the standard prior.