Institution profile

RAND Corporation

Academic institutionnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Designing Incident Reporting Systems for Harms from General-Purpose AI

Nov 08, 2025

In response to escalating safety and rights risks posed by general-purpose artificial intelligence (GPAI), this paper proposes the first systematic reporting framework for GPAI incidents. Drawing on a systematic literature review and cross-case analysis of high-stakes domains—including aviation and healthcare—as well as regulatory practices in the U.S. and EU, the study identifies seven core dimensions: policy objectives, reporting entities, incident typologies, reporting modalities (mandatory vs. voluntary), near-miss inclusion, anonymity safeguards, and legal immunity provisions. It critically examines the trade-offs among safety learning, cross-organizational information sharing, and legal interoperability inherent in each mechanism. The resulting framework offers policymakers and researchers an actionable, theory-informed blueprint for designing GPAI incident reporting infrastructure—addressing a critical gap in GPAI risk governance and advancing the institutional foundations for responsible AI development and deployment.

1 citationsRead paper

Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

Mar 05, 2026

This work addresses the growing reliance on large language models (LLMs) as automated judges in AI evaluation, despite a lack of systematic validation of their reliability. We propose the first open-source stress-testing framework specifically designed for LLM-based judges, which automatically generates perturbations in text formatting, phrasing, and level of detail to assess accuracy and robustness in both binary classification and ordinal scoring tasks. The framework incorporates multidimensional metrics and spans diverse scenarios, including free-form responses and agent-based tasks. Experiments with four state-of-the-art LLM judges across four established benchmarks reveal that none consistently maintain reliable performance across all settings, exposing significant robustness deficiencies in current judge systems.

0 citationsRead paper

Small models, big threats: Characterizing safety challenges from low-compute AI models

Jan 29, 2026

This study addresses a critical gap in AI governance by demonstrating that the prevailing focus on high-compute models overlooks emerging safety risks from low-compute models, whose capabilities have surged due to algorithmic advances. Through systematic analysis of over 5,000 open-source large language models, combined with historical benchmarking, parameter quantization, resource simulation, and deployment experiments on consumer-grade hardware, we show that model compression and agent-based workflows are enabling hazardous capabilities—such as disinformation generation and voice-cloning fraud—to migrate efficiently to lightweight models. Our findings reveal that the model scale required to match mainstream LLM performance has decreased by more than an order of magnitude within a year, rendering many digital-society harms executable on everyday devices. This challenges compute-centric regulatory paradigms and underscores the urgent need for governance frameworks that account for low-compute AI risks.

0 citationsRead paper

Estimating the Impact of Case Management in MDLs: Lone Pine Orders and Bellwether Trials

Dec 08, 2025

This study examines how judicial case management affects outcomes in multidistrict litigation (MDL), focusing on Lone Pine orders and bellwether trials as mechanisms to mitigate coercive settlement pressures arising from information asymmetry and non-meritorious claims. Leveraging a comprehensive, longitudinal MDL panel dataset spanning 1992–2017, we employ an event-study design combined with a difference-in-differences (DID) strategy—integrating judicial opinions and case disposition records—to causally identify the impact of Lone Pine orders on MDL resolution rates. Results show that Lone Pine orders significantly increase the number of cases resolved within MDL proceedings, enhancing both procedural efficiency and claim-screening quality. Our contribution is the first nationally representative, long-term, causally identified empirical evidence on judicial management tools in MDLs, providing critical policy insights for improving class-action governance and federal judicial administration.

0 citationsRead paper
Recent publications

Latest Papers

Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

Mar 05, 2026

This work addresses the growing reliance on large language models (LLMs) as automated judges in AI evaluation, despite a lack of systematic validation of their reliability. We propose the first open-source stress-testing framework specifically designed for LLM-based judges, which automatically generates perturbations in text formatting, phrasing, and level of detail to assess accuracy and robustness in both binary classification and ordinal scoring tasks. The framework incorporates multidimensional metrics and spans diverse scenarios, including free-form responses and agent-based tasks. Experiments with four state-of-the-art LLM judges across four established benchmarks reveal that none consistently maintain reliable performance across all settings, exposing significant robustness deficiencies in current judge systems.

0 citationsRead paper

Small models, big threats: Characterizing safety challenges from low-compute AI models

Jan 29, 2026

This study addresses a critical gap in AI governance by demonstrating that the prevailing focus on high-compute models overlooks emerging safety risks from low-compute models, whose capabilities have surged due to algorithmic advances. Through systematic analysis of over 5,000 open-source large language models, combined with historical benchmarking, parameter quantization, resource simulation, and deployment experiments on consumer-grade hardware, we show that model compression and agent-based workflows are enabling hazardous capabilities—such as disinformation generation and voice-cloning fraud—to migrate efficiently to lightweight models. Our findings reveal that the model scale required to match mainstream LLM performance has decreased by more than an order of magnitude within a year, rendering many digital-society harms executable on everyday devices. This challenges compute-centric regulatory paradigms and underscores the urgent need for governance frameworks that account for low-compute AI risks.

0 citationsRead paper

Estimating the Impact of Case Management in MDLs: Lone Pine Orders and Bellwether Trials

Dec 08, 2025

This study examines how judicial case management affects outcomes in multidistrict litigation (MDL), focusing on Lone Pine orders and bellwether trials as mechanisms to mitigate coercive settlement pressures arising from information asymmetry and non-meritorious claims. Leveraging a comprehensive, longitudinal MDL panel dataset spanning 1992–2017, we employ an event-study design combined with a difference-in-differences (DID) strategy—integrating judicial opinions and case disposition records—to causally identify the impact of Lone Pine orders on MDL resolution rates. Results show that Lone Pine orders significantly increase the number of cases resolved within MDL proceedings, enhancing both procedural efficiency and claim-screening quality. Our contribution is the first nationally representative, long-term, causally identified empirical evidence on judicial management tools in MDLs, providing critical policy insights for improving class-action governance and federal judicial administration.

0 citationsRead paper

Designing Incident Reporting Systems for Harms from General-Purpose AI

Nov 08, 2025

In response to escalating safety and rights risks posed by general-purpose artificial intelligence (GPAI), this paper proposes the first systematic reporting framework for GPAI incidents. Drawing on a systematic literature review and cross-case analysis of high-stakes domains—including aviation and healthcare—as well as regulatory practices in the U.S. and EU, the study identifies seven core dimensions: policy objectives, reporting entities, incident typologies, reporting modalities (mandatory vs. voluntary), near-miss inclusion, anonymity safeguards, and legal immunity provisions. It critically examines the trade-offs among safety learning, cross-organizational information sharing, and legal interoperability inherent in each mechanism. The resulting framework offers policymakers and researchers an actionable, theory-informed blueprint for designing GPAI incident reporting infrastructure—addressing a critical gap in GPAI risk governance and advancing the institutional foundations for responsible AI development and deployment.

1 citationsRead paper