Institution profile

Model Evaluation & Threat Research

Research institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology

Feb 18, 2026

This study evaluates whether medium-to-large language models (LLMs) can enhance the ability of novice biologists to perform viral reverse genetics in a real-world wet-lab setting, using a preregistered, investigator-blinded, randomized controlled trial (June–August 2025; n=153). While LLM assistance did not significantly improve overall protocol completion rates (5.2% vs. 6.6%, P=0.759), it yielded better performance in four of five subtasks, notably in cell culture (68.8% vs. 55.3%, P=0.059). Bayesian analysis estimated a ~1.4-fold increase in typical task success and a higher likelihood of advancing through intermediate steps (posterior probabilities: 81%–96%). This work provides the first empirical quantification of LLMs’ impact on hands-on biological experimentation, revealing a performance gap between in silico benchmarks and real-world application, and offering critical evidence for AI-driven biosafety risk assessment.

0 citationsRead paper

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Jul 11, 2025

This study investigates the causal impact of AI programming tools on the productivity of experienced open-source developers in early 2025. Method: We conduct the first randomized controlled trial (RCT) embedded in mature, real-world open-source projects, comparing an intervention group using Cursor Pro with Claude 3.5/3.7 Sonnet against a control group. Primary outcome is task completion time; secondary measures include developer self-assessments and expert predictions for validation. Contribution/Results: Contrary to developers’ expectation of a 24% speedup, AI assistance significantly *increased* mean task time by 19% in high-maturity projects—a negative effect robust across multiple sensitivity and specification tests. This is the first RCT to empirically reveal a counterintuitive inverse relationship between AI assistance and expert developer efficiency in authentic open-source settings. We systematically identify and evaluate 20 potential moderating factors, providing critical empirical evidence on the boundary conditions under which AI augments—or impedes—software development productivity.

0 citationsRead paper
Recent publications

Latest Papers

Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology

Feb 18, 2026

This study evaluates whether medium-to-large language models (LLMs) can enhance the ability of novice biologists to perform viral reverse genetics in a real-world wet-lab setting, using a preregistered, investigator-blinded, randomized controlled trial (June–August 2025; n=153). While LLM assistance did not significantly improve overall protocol completion rates (5.2% vs. 6.6%, P=0.759), it yielded better performance in four of five subtasks, notably in cell culture (68.8% vs. 55.3%, P=0.059). Bayesian analysis estimated a ~1.4-fold increase in typical task success and a higher likelihood of advancing through intermediate steps (posterior probabilities: 81%–96%). This work provides the first empirical quantification of LLMs’ impact on hands-on biological experimentation, revealing a performance gap between in silico benchmarks and real-world application, and offering critical evidence for AI-driven biosafety risk assessment.

0 citationsRead paper

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Jul 11, 2025

This study investigates the causal impact of AI programming tools on the productivity of experienced open-source developers in early 2025. Method: We conduct the first randomized controlled trial (RCT) embedded in mature, real-world open-source projects, comparing an intervention group using Cursor Pro with Claude 3.5/3.7 Sonnet against a control group. Primary outcome is task completion time; secondary measures include developer self-assessments and expert predictions for validation. Contribution/Results: Contrary to developers’ expectation of a 24% speedup, AI assistance significantly *increased* mean task time by 19% in high-maturity projects—a negative effect robust across multiple sensitivity and specification tests. This is the first RCT to empirically reveal a counterintuitive inverse relationship between AI assistance and expert developer efficiency in authentic open-source settings. We systematically identify and evaluate 20 potential moderating factors, providing critical empirical evidence on the boundary conditions under which AI augments—or impedes—software development productivity.

0 citationsRead paper