🤖 AI Summary
Parameter estimation under missing data is doubly sensitive to both misspecification of the underlying data model and deviations from standard missingness mechanisms (MCAR/MAR/MNAR) or Huber-type contamination.
Method: This paper proposes a robust M-estimation framework grounded in the Maximum Mean Discrepancy (MMD), leveraging kernel embeddings and functional analysis tools. It avoids explicit specification of either the missingness mechanism or the full-data distribution.
Contribution/Results: The method achieves joint robustness against both types of misspecification—theoretically guaranteeing strong consistency and asymptotic normality under MCAR. It yields a decomposable, explicit error bound that cleanly separates model misspecification error from missingness-induced bias. Moreover, it maintains controlled estimation error under MNAR and Huber contamination. By circumventing stringent modeling assumptions, the approach significantly enhances robustness and reliability in practical applications.
📝 Abstract
In the missing data literature, the Maximum Likelihood Estimator (MLE) is celebrated for its ignorability property under missing at random (MAR) data. However, its sensitivity to misspecification of the (complete) data model, even under MAR, remains a significant limitation. This issue is further exacerbated by the fact that the MAR assumption may not always be realistic, introducing an additional source of potential misspecification through the missingness mechanism. To address this, we propose a novel M-estimation procedure based on the Maximum Mean Discrepancy (MMD), which is provably robust to both model misspecification and deviations from the assumed missingness mechanism. Our approach offers strong theoretical guarantees and improved reliability in complex settings. We establish the consistency and asymptotic normality of the estimator under missingness completely at random (MCAR), provide an efficient stochastic gradient descent algorithm, and derive error bounds that explicitly separate the contributions of model misspecification and missingness bias. Furthermore, we analyze missing not at random (MNAR) scenarios where our estimator maintains controlled error, including a Huber setting where both the missingness mechanism and the data model are contaminated. Our contributions refine the understanding of the limitations of the MLE and provide a robust and principled alternative for handling missing data.