Institution profile

FutureHouse

Research institution
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Training a Scientific Reasoning Model for Chemistry

Jun 04, 2025arXiv.org

Existing chemical language models typically require domain-specific pretraining, limiting data efficiency and generalizability in reasoning across diverse experimental tasks. Method: We propose a novel paradigm for building high-performance chemical reasoning models via post-training only—eliminating the need for domain-specific pretraining. Leveraging the Mistral-Small-24B architecture, we apply reinforcement learning–based chain-of-thought fine-tuning on over 640,000 experimentally annotated chemistry problems, enabling joint natural-language and SMILES-based structural reasoning across 375 experiment-driven tasks—including synthetic feasibility, pharmacokinetics, receptor activity, and odor prediction. Contribution/Results: This work achieves, for the first time, zero-domain-pretraining chemical reasoning modeling. Our data efficiency exceeds that of specialized models by over one order of magnitude. The resulting model, ether0, outperforms state-of-the-art general-purpose and multimodal chemical foundation models—and even human experts—on molecular design benchmarks.

13 citations2 influentialRead paper

BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology

Feb 28, 2025

A comprehensive benchmark for evaluating end-to-end scientific reasoning capabilities of LLM agents in bioinformatics is lacking. Method: We introduce BixBench, the first open-source benchmark tailored to computational biology, comprising 50+ realistic analytical scenarios and nearly 300 open-ended questions. It systematically assesses agent performance across data exploration, multi-step tool invocation, dynamic planning, and complex result interpretation. We explicitly define and quantify both reasoning and execution fidelity of LLM agents on authentic bioinformatic tasks. Contribution/Results: Our evaluation reveals critical bottlenecks: state-of-the-art models achieve only 17% accuracy on open-ended questions and perform no better than random chance on multiple-choice items. We release an open-source agent framework built on GPT-4o and Claude 3.5 Sonnet, enabling reproducible, multi-step analytical assessment—filling a key methodological gap and advancing rigorous evaluation and development of biological AI agents.

0 citationsRead paper
Recent publications

Latest Papers

Training a Scientific Reasoning Model for Chemistry

Jun 04, 2025arXiv.org

Existing chemical language models typically require domain-specific pretraining, limiting data efficiency and generalizability in reasoning across diverse experimental tasks. Method: We propose a novel paradigm for building high-performance chemical reasoning models via post-training only—eliminating the need for domain-specific pretraining. Leveraging the Mistral-Small-24B architecture, we apply reinforcement learning–based chain-of-thought fine-tuning on over 640,000 experimentally annotated chemistry problems, enabling joint natural-language and SMILES-based structural reasoning across 375 experiment-driven tasks—including synthetic feasibility, pharmacokinetics, receptor activity, and odor prediction. Contribution/Results: This work achieves, for the first time, zero-domain-pretraining chemical reasoning modeling. Our data efficiency exceeds that of specialized models by over one order of magnitude. The resulting model, ether0, outperforms state-of-the-art general-purpose and multimodal chemical foundation models—and even human experts—on molecular design benchmarks.

13 citations2 influentialRead paper

BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology

Feb 28, 2025

A comprehensive benchmark for evaluating end-to-end scientific reasoning capabilities of LLM agents in bioinformatics is lacking. Method: We introduce BixBench, the first open-source benchmark tailored to computational biology, comprising 50+ realistic analytical scenarios and nearly 300 open-ended questions. It systematically assesses agent performance across data exploration, multi-step tool invocation, dynamic planning, and complex result interpretation. We explicitly define and quantify both reasoning and execution fidelity of LLM agents on authentic bioinformatic tasks. Contribution/Results: Our evaluation reveals critical bottlenecks: state-of-the-art models achieve only 17% accuracy on open-ended questions and perform no better than random chance on multiple-choice items. We release an open-source agent framework built on GPT-4o and Claude 3.5 Sonnet, enabling reproducible, multi-step analytical assessment—filling a key methodological gap and advancing rigorous evaluation and development of biological AI agents.

0 citationsRead paper