Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
This study evaluates whether medium-to-large language models (LLMs) can enhance the ability of novice biologists to perform viral reverse genetics in a real-world wet-lab setting, using a preregistered, investigator-blinded, randomized controlled trial (June–August 2025; n=153). While LLM assistance did not significantly improve overall protocol completion rates (5.2% vs. 6.6%, P=0.759), it yielded better performance in four of five subtasks, notably in cell culture (68.8% vs. 55.3%, P=0.059). Bayesian analysis estimated a ~1.4-fold increase in typical task success and a higher likelihood of advancing through intermediate steps (posterior probabilities: 81%–96%). This work provides the first empirical quantification of LLMs’ impact on hands-on biological experimentation, revealing a performance gap between in silico benchmarks and real-world application, and offering critical evidence for AI-driven biosafety risk assessment.