Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
This work addresses the limitations of existing medical reasoning verification methods, which rely on single-step retrieval and provide only scalar rewards, thereby lacking interpretability and dynamic knowledge acquisition. The authors propose a tool-augmented reinforcement learning agent framework that iteratively queries external medical corpora during verification and integrates trajectory-supervised iterative reinforcement learning with an adaptive curriculum mechanism. This approach enables dynamic evidence retrieval for the first time, substantially improving both verification efficiency and reliability. Evaluated on four medical reasoning benchmarks, the method significantly outperforms current state-of-the-art approaches, achieving a 23.5% absolute accuracy gain on MedQA and a 32.0% improvement on MedXpertQA, while reducing the required sampling budget by a factor of eight.