DeepImagine: Learning Biomedical Reasoning via Successive Counterfactual Imagining
Current large language models struggle to capture the underlying causal mechanisms when predicting clinical trial outcomes. To address this limitation, this work proposes a novel training paradigm that integrates supervised fine-tuning, reinforcement learning guided by validation-based rewards, and synthetic counterfactual reasoning trajectories. By leveraging counterfactual pairs to construct training data, the approach steers models—particularly those under 10B parameters, such as Qwen3.5-9B—toward learning interpretable biomedical causal reasoning processes through iterative counterfactual imagination. The method substantially outperforms both unadapted language models and conventional correlation-based baselines, while simultaneously generating transparent and human-interpretable reasoning pathways that reflect plausible causal structures in clinical contexts.