π€ AI Summary
This study addresses non-random missingness in millions of traffic stop records from the Stanford Open Policing Project, particularly widespread unrecorded race information. Method: We propose the first systematic missingness-sensitivity analysis framework for policing data, integrating multi-metric missingness-mechanism tests, missing-pattern visualization, hypothesis-driven bias correction, and an extension of sharp bounds to robustly estimate average treatment effects. Contribution/Results: We provide the first quantitative evidence that missing race data severely biases causal inferencesβe.g., inflating or reversing estimated racial disparities in stop rates. Our framework establishes a transferable paradigm for missingness sensitivity analysis, offering a generalizable solution to non-random missingness in fairness-aware social science research and policy evaluation. The approach is computationally efficient, interpretable, and grounded in formal identification theory, enabling practitioners to assess the robustness of conclusions to plausible missingness assumptions without requiring strong ignorability assumptions.
π Abstract
In this article we explore the data available through the Stanford Open Policing Project. The data consist of information on millions of traffic stops across close to 100 different cities and highway patrols. Using a variety of metrics, we identify that the data is not missing completely at random. Furthermore, we develop ways of quantifying and visualizing missingness trends for different variables across the datasets. We follow up by performing a sensitivity analysis to extend work done on the outcome test as well as to extend work done on sharp bounds on the average treatment effect. We demonstrate that bias calculations can fundamentally shift depending on the assumptions made about the observations for which the race variable has not been recorded. We suggest ways that our missingness sensitivity analysis can be extended to myriad different contexts.