Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
This study addresses critical limitations in evaluating medical large language models (LLMs)—namely benchmark saturation, closed data access, and insufficient task coverage—by introducing the first fully open-source evaluation suite encompassing 30 diverse tasks, including question answering, information extraction, medical calculation, and open-ended clinical reasoning. The authors systematically assess 61 models across 71 configurations using an automated, multi-model, multi-task framework that integrates verifiable metrics with LLM-as-a-Judge methodologies; selected subsets also function as reinforcement learning environments to enhance medical reasoning capabilities. Experimental results reveal that state-of-the-art reasoning models achieve the strongest performance, domain-specific fine-tuned models significantly outperform general-purpose counterparts, closed-source models exhibit greater token efficiency, and most models display answer-order bias.