PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of expensive manual annotation in post-training large models for biological reasoning by proposing a novel paradigm that leverages cellular perturbation atlases to construct reinforcement learning environments. Through credible trajectory-supervised initialization, gene-pathway multi-scale rewards, and forward perturbation prediction, we achieve autonomous training driven entirely by experimental data. This approach generates high-quality multi-scale biological representations without additional fine-tuning, significantly enhancing reasoning capabilities in unseen scenarios. Furthermore, it demonstrates successful zero-shot transferability across diverse downstream tasks, validating both the feasibility and scalability of utilizing large-scale experimental data to drive general-purpose biological reasoning.
📝 Abstract
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Trained only on forward perturbation-response prediction, PertMind improved response inference in unseen cellular contexts while retaining general language capabilities. It also transferred without task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generated biological profiles that supported competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.
Problem

Research questions and friction points this paper is trying to address.

Biological Reasoning
Large Language Models
Post-training Scalability
Cellular Perturbation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Cellular Perturbation
Biological Reasoning
PertMind
Zero-shot Transfer
🔎 Similar Papers
No similar papers found.