Institution profile

Insilico Medicine

Industry researchnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Jul 20, 2026

This study investigates whether general-purpose large language models (LLMs) can effectively handle complex three-dimensional spatial constraints—such as protein binding pockets, pharmacophores, and anchor fragments—in structure-based drug design to generate plausible binding molecules. To this end, the authors introduce 3D-Fit, the first evaluation framework tailored for multi-spatial-constraint scenarios, which integrates a token-efficient benchmarking strategy with a constraint modeling approach conditioned on heterogeneous spatial information. Experimental results demonstrate that, although current LLMs slightly underperform state-of-the-art diffusion models in generation quality, they already exhibit a promising capacity to jointly incorporate diverse 3D constraints and show strong potential for scalability, thereby validating their prospective utility in structure-based drug design.

0 citationsRead paper

URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

Jul 06, 2026

Current retrosynthetic planning systems lack a flexible and chemically interpretable evaluation benchmark, making objective performance comparisons difficult. To address this gap, this work proposes the URSA evaluation framework, which introduces— for the first time—a dimension of chemical plausibility that simulates expert chemists’ judgment, complementing formal correctness. URSA enables a comprehensive assessment of retrosynthetic routes by considering reaction feasibility, pathway convergence, and other chemically relevant criteria. The framework evaluates both specialized models and large language models in realistic drug design scenarios, revealing that while the latter show promise in high-level strategic planning, they still significantly lag behind specialized models in task reliability. URSA thus establishes a new, interpretable, and holistic benchmark for the practical evaluation of retrosynthetic methodologies.

0 citationsRead paper

MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery

Mar 03, 2026

This work addresses the limited scientific reasoning capabilities of general-purpose large language models (LLMs) in drug discovery, which cannot be reliably improved merely by scaling up model size or incorporating generic reasoning mechanisms. To overcome this, the authors introduce MMAI Gym for Science—a platform that integrates multimodal molecular data with task-specific training, inference, and evaluation frameworks—and propose a lightweight Liquid Foundation Model architecture designed to effectively model “molecular language.” This approach achieves state-of-the-art or competitive performance across key tasks including molecular optimization, ADMET prediction, retrosynthesis, and drug–target activity prediction, while requiring substantially lower computational resources than both larger general-purpose and specialized models, thereby breaking the performance bottleneck of general LLMs in specialized scientific domains.

0 citationsRead paper

FLANS at SemEval-2026 Task 7: RAG with Open-Sourced Smaller LLMs for Everyday Knowledge Across Diverse Languages and Cultures

Mar 02, 2026

This work addresses the challenge of everyday knowledge question answering—including both short-answer and multiple-choice formats—in multilingual and cross-cultural settings. The authors propose a retrieval-augmented generation (RAG) framework that integrates open-source small language models deployed via Ollama, a custom culture-aware knowledge base (CulKBs) combining localized Wikipedia content with real-time DuckDuckGo search results, and iterative prompt engineering. Evaluated on SemEval-2026 Task 7, the system supports English, Spanish, and Chinese while prioritizing user privacy and computational sustainability. All code and resources are publicly released to advance research on lightweight, culturally adaptive multilingual question answering systems.

0 citationsRead paper
Recent publications

Latest Papers

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Jul 20, 2026

This study investigates whether general-purpose large language models (LLMs) can effectively handle complex three-dimensional spatial constraints—such as protein binding pockets, pharmacophores, and anchor fragments—in structure-based drug design to generate plausible binding molecules. To this end, the authors introduce 3D-Fit, the first evaluation framework tailored for multi-spatial-constraint scenarios, which integrates a token-efficient benchmarking strategy with a constraint modeling approach conditioned on heterogeneous spatial information. Experimental results demonstrate that, although current LLMs slightly underperform state-of-the-art diffusion models in generation quality, they already exhibit a promising capacity to jointly incorporate diverse 3D constraints and show strong potential for scalability, thereby validating their prospective utility in structure-based drug design.

0 citationsRead paper

URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

Jul 06, 2026

Current retrosynthetic planning systems lack a flexible and chemically interpretable evaluation benchmark, making objective performance comparisons difficult. To address this gap, this work proposes the URSA evaluation framework, which introduces— for the first time—a dimension of chemical plausibility that simulates expert chemists’ judgment, complementing formal correctness. URSA enables a comprehensive assessment of retrosynthetic routes by considering reaction feasibility, pathway convergence, and other chemically relevant criteria. The framework evaluates both specialized models and large language models in realistic drug design scenarios, revealing that while the latter show promise in high-level strategic planning, they still significantly lag behind specialized models in task reliability. URSA thus establishes a new, interpretable, and holistic benchmark for the practical evaluation of retrosynthetic methodologies.

0 citationsRead paper

MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery

Mar 03, 2026

This work addresses the limited scientific reasoning capabilities of general-purpose large language models (LLMs) in drug discovery, which cannot be reliably improved merely by scaling up model size or incorporating generic reasoning mechanisms. To overcome this, the authors introduce MMAI Gym for Science—a platform that integrates multimodal molecular data with task-specific training, inference, and evaluation frameworks—and propose a lightweight Liquid Foundation Model architecture designed to effectively model “molecular language.” This approach achieves state-of-the-art or competitive performance across key tasks including molecular optimization, ADMET prediction, retrosynthesis, and drug–target activity prediction, while requiring substantially lower computational resources than both larger general-purpose and specialized models, thereby breaking the performance bottleneck of general LLMs in specialized scientific domains.

0 citationsRead paper

FLANS at SemEval-2026 Task 7: RAG with Open-Sourced Smaller LLMs for Everyday Knowledge Across Diverse Languages and Cultures

Mar 02, 2026

This work addresses the challenge of everyday knowledge question answering—including both short-answer and multiple-choice formats—in multilingual and cross-cultural settings. The authors propose a retrieval-augmented generation (RAG) framework that integrates open-source small language models deployed via Ollama, a custom culture-aware knowledge base (CulKBs) combining localized Wikipedia content with real-time DuckDuckGo search results, and iterative prompt engineering. Evaluated on SemEval-2026 Task 7, the system supports English, Spanish, and Chinese while prioritizing user privacy and computational sustainability. All code and resources are publicly released to advance research on lightweight, culturally adaptive multilingual question answering systems.

0 citationsRead paper