Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Assessing large language models’ (LLMs) capability to debug high-consequence virology experimental protocols—and their potential role in dual-use technology governance—remains unexplored. Method: We introduce VCT, the first multimodal question-answering benchmark for virology, comprising 322 expert-crafted questions spanning foundational, tacit, and visual knowledge domains; we further propose a multimodal LLM evaluation framework integrating text, flowchart, and experimental image understanding. Contribution/Results: This work presents the first systematic quantification of LLM performance on dual-use virological procedural reasoning: the o3 model achieves 43.8% accuracy—significantly exceeding both the human expert average (22.1%) and 94% of individual experts. These findings provide empirical evidence and a methodological foundation for integrating LLMs into life science dual-use governance frameworks.