Institution profile

MedARC

Academic institutionnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

May 02, 2026

This study addresses critical limitations in evaluating medical large language models (LLMs)—namely benchmark saturation, closed data access, and insufficient task coverage—by introducing the first fully open-source evaluation suite encompassing 30 diverse tasks, including question answering, information extraction, medical calculation, and open-ended clinical reasoning. The authors systematically assess 61 models across 71 configurations using an automated, multi-model, multi-task framework that integrates verifiable metrics with LLM-as-a-Judge methodologies; selected subsets also function as reinforcement learning environments to enhance medical reasoning capabilities. Experimental results reveal that state-of-the-art reasoning models achieve the strongest performance, domain-specific fine-tuned models significantly outperform general-purpose counterparts, closed-source models exhibit greater token efficiency, and most models display answer-order bias.

0 citationsRead paper
Recent publications

Latest Papers

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

May 02, 2026

This study addresses critical limitations in evaluating medical large language models (LLMs)—namely benchmark saturation, closed data access, and insufficient task coverage—by introducing the first fully open-source evaluation suite encompassing 30 diverse tasks, including question answering, information extraction, medical calculation, and open-ended clinical reasoning. The authors systematically assess 61 models across 71 configurations using an automated, multi-model, multi-task framework that integrates verifiable metrics with LLM-as-a-Judge methodologies; selected subsets also function as reinforcement learning environments to enhance medical reasoning capabilities. Experimental results reveal that state-of-the-art reasoning models achieve the strongest performance, domain-specific fine-tuned models significantly outperform general-purpose counterparts, closed-source models exhibit greater token efficiency, and most models display answer-order bias.

0 citationsRead paper