Training a Scientific Reasoning Model for Chemistry
Existing chemical language models typically require domain-specific pretraining, limiting data efficiency and generalizability in reasoning across diverse experimental tasks. Method: We propose a novel paradigm for building high-performance chemical reasoning models via post-training only—eliminating the need for domain-specific pretraining. Leveraging the Mistral-Small-24B architecture, we apply reinforcement learning–based chain-of-thought fine-tuning on over 640,000 experimentally annotated chemistry problems, enabling joint natural-language and SMILES-based structural reasoning across 375 experiment-driven tasks—including synthetic feasibility, pharmacokinetics, receptor activity, and odor prediction. Contribution/Results: This work achieves, for the first time, zero-domain-pretraining chemical reasoning modeling. Our data efficiency exceeds that of specialized models by over one order of magnitude. The resulting model, ether0, outperforms state-of-the-art general-purpose and multimodal chemical foundation models—and even human experts—on molecular design benchmarks.