Institution profile

SecureBio

Industry researchnorthamerica · us
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

Sep 05, 2026

We introduce ABLE, a benchmark for evaluating LLM agents'ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.

5 citationsRead paper

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

Jul 15, 2026

This work addresses critical safety challenges faced by large language models in the biological domain, where models either risk disclosing hazardous information or overly restrict legitimate scientific inquiries due to a lack of precise risk identification and differentiated control mechanisms. To resolve this, the authors propose BioTIER—the first risk-tiered biosafety evaluation benchmark—which categorizes biological content into three distinct classes: catastrophe-averting, dual-use research of concern, and general biology. Leveraging an expert-curated dataset of 542 metadata-annotated prompts, BioTIER enables accurate discrimination between high-risk information and beneficial scientific knowledge. The benchmark facilitates targeted refusal strategies that effectively block a minimal set of catastrophic content while preserving open access to the vast majority of research-relevant knowledge, thereby significantly enhancing both model safety and scientific utility.

0 citationsRead paper

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Jun 09, 2026

This study systematically evaluates the capabilities of large language model (LLM) agents in biosafety-relevant tasks, balancing their scientific potential against misuse risks. We introduce the first benchmark suite encompassing both benign and dual-use biological tasks, requiring models to integrate biological knowledge with programming skills to perform high-risk operations such as liquid-handling robot control, DNA fragment design, and circumvention of synthetic screening protocols. Crucially, we incorporate wet-lab validation to assess real-world feasibility. Results demonstrate that all evaluated models surpass the median performance of human experts, with scripts generated by o4-mini-high successfully executing DNA assembly on a physical robotic platform—providing the first empirical evidence of end-to-end LLM agent feasibility in complex biological experimentation.

0 citationsRead paper

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

Aug 13, 2025

Current AI model evaluations in chemical and biological (ChemBio) safety suffer from opaque reporting and a lack of standardized disclosure practices. Method: This paper introduces the first transparent reporting standard specifically for ChemBio risk assessment of AI models. Drawing on best practices from government, academia, and industry, we develop a structured reporting framework, a standardized evaluation metadata schema, a concise three-page operational report template, and multiple “gold-standard” exemplar reports. Contribution/Results: We systematically define, for the first time, the disclosure dimensions and quality requirements for ChemBio safety evaluations; significantly improve the completeness and reproducibility of assessment information; and enable third-party independent auditing and cross-model comparability. The standard has been adopted by multiple leading AI research organizations, enhancing both public trust and methodological rigor in ChemBio safety assessment.

0 citationsRead paper
Recent publications

Latest Papers

Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

Sep 05, 2026

We introduce ABLE, a benchmark for evaluating LLM agents'ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.

5 citationsRead paper

BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation

Jul 15, 2026

This work addresses critical safety challenges faced by large language models in the biological domain, where models either risk disclosing hazardous information or overly restrict legitimate scientific inquiries due to a lack of precise risk identification and differentiated control mechanisms. To resolve this, the authors propose BioTIER—the first risk-tiered biosafety evaluation benchmark—which categorizes biological content into three distinct classes: catastrophe-averting, dual-use research of concern, and general biology. Leveraging an expert-curated dataset of 542 metadata-annotated prompts, BioTIER enables accurate discrimination between high-risk information and beneficial scientific knowledge. The benchmark facilitates targeted refusal strategies that effectively block a minimal set of catastrophic content while preserving open access to the vast majority of research-relevant knowledge, thereby significantly enhancing both model safety and scientific utility.

0 citationsRead paper

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Jun 09, 2026

This study systematically evaluates the capabilities of large language model (LLM) agents in biosafety-relevant tasks, balancing their scientific potential against misuse risks. We introduce the first benchmark suite encompassing both benign and dual-use biological tasks, requiring models to integrate biological knowledge with programming skills to perform high-risk operations such as liquid-handling robot control, DNA fragment design, and circumvention of synthetic screening protocols. Crucially, we incorporate wet-lab validation to assess real-world feasibility. Results demonstrate that all evaluated models surpass the median performance of human experts, with scripts generated by o4-mini-high successfully executing DNA assembly on a physical robotic platform—providing the first empirical evidence of end-to-end LLM agent feasibility in complex biological experimentation.

0 citationsRead paper

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

Aug 13, 2025

Current AI model evaluations in chemical and biological (ChemBio) safety suffer from opaque reporting and a lack of standardized disclosure practices. Method: This paper introduces the first transparent reporting standard specifically for ChemBio risk assessment of AI models. Drawing on best practices from government, academia, and industry, we develop a structured reporting framework, a standardized evaluation metadata schema, a concise three-page operational report template, and multiple “gold-standard” exemplar reports. Contribution/Results: We systematically define, for the first time, the disclosure dimensions and quality requirements for ChemBio safety evaluations; significantly improve the completeness and reproducibility of assessment information; and enable third-party independent auditing and cross-model comparability. The standard has been adopted by multiple leading AI research organizations, enhancing both public trust and methodological rigor in ChemBio safety assessment.

0 citationsRead paper