Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology

πŸ“… 2026-02-18
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study evaluates whether medium-to-large language models (LLMs) can enhance the ability of novice biologists to perform viral reverse genetics in a real-world wet-lab setting, using a preregistered, investigator-blinded, randomized controlled trial (June–August 2025; n=153). While LLM assistance did not significantly improve overall protocol completion rates (5.2% vs. 6.6%, P=0.759), it yielded better performance in four of five subtasks, notably in cell culture (68.8% vs. 55.3%, P=0.059). Bayesian analysis estimated a ~1.4-fold increase in typical task success and a higher likelihood of advancing through intermediate steps (posterior probabilities: 81%–96%). This work provides the first empirical quantification of LLMs’ impact on hands-on biological experimentation, revealing a performance gap between in silico benchmarks and real-world application, and offering critical evidence for AI-driven biosafety risk assessment.

Technology Category

Application Category

πŸ“ Abstract
Large language models (LLMs) perform strongly on biological benchmarks, raising concerns that they may help novice actors acquire dual-use laboratory skills. Yet, whether this translates to improved human performance in the physical laboratory remains unclear. To address this, we conducted a pre-registered, investigator-blinded, randomized controlled trial (June-August 2025; n = 153) evaluating whether LLMs improve novice performance in tasks that collectively model a viral reverse genetics workflow. We observed no significant difference in the primary endpoint of workflow completion (5.2% LLM vs. 6.6% Internet; P = 0.759), nor in the success rate of individual tasks. However, the LLM arm had numerically higher success rates in four of the five tasks, most notably for the cell culture task (68.8% LLM vs. 55.3% Internet; P = 0.059). Post-hoc Bayesian modeling of pooled data estimates an approximate 1.4-fold increase (95% CrI 0.74-2.62) in success for a "typical" reverse genetics task under LLM assistance. Ordinal regression modelling suggests that participants in the LLM arm were more likely to progress through intermediate steps across all tasks (posterior probability of a positive effect: 81%-96%). Overall, mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit. These results reveal a gap between in silico benchmarks and real-world utility, underscoring the need for physical-world validation of AI biosecurity assessments as model capabilities and user proficiency evolve.
Problem

Research questions and friction points this paper is trying to address.

LLM-assisted learning
novice performance
biological laboratory tasks
reverse genetics workflow
AI biosecurity
Innovation

Methods, ideas, or system contributions that make the work stand out.

randomized controlled trial
large language models
laboratory performance
biosecurity
physical-world validation
πŸ”Ž Similar Papers
No similar papers found.
S
Shen Zhou Hong
Active Site, Cambridge, MA 02142, United States
A
Alex Kleinman
Active Site, Cambridge, MA 02142, United States
A
Alyssa Mathiowetz
Active Site, Cambridge, MA 02142, United States
A
Adam Howes
Independent
J
Julian Cohen
Active Site, Cambridge, MA 02142, United States
S
Suveer Ganta
Active Site, Cambridge, MA 02142, United States
A
Alex Letizia
Active Site, Cambridge, MA 02142, United States
D
Dora Liao
Active Site, Cambridge, MA 02142, United States
D
Deepika Pahari
Active Site, Cambridge, MA 02142, United States
X
Xavier Roberts-Gaal
Active Site, Cambridge, MA 02142, United States
L
Luca Righetti
Model Evaluation and Threat Research, Inc., Covina, CA 91723, United States
J
Joe Torres
Active Site, Cambridge, MA 02142, United States