Institution profile

Medical Centre of Postgraduate Education

Academic institutioneurope · pl
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Jun 10, 2026

This study addresses the susceptibility of existing multiple-choice question answering (MCQA)–based evaluations of medical large language models to guessing and answer bias, which often inflate estimates of true clinical reasoning capabilities. To mitigate these limitations, the authors introduce a more rigorous benchmark based on Polish medical licensing examinations, incorporating over 15,000 new questions spanning two additional clinical domains and implementing four structural enhancements designed to reduce inherent MCQA biases. The benchmark facilitates cross-lingual evaluation and data contamination detection, and was used to systematically assess 21 prominent large language models. Results reveal that under this more challenging setting, the top-performing model, Qwen3.5-122B, exhibits performance drops of 28.4 and 31 percentage points on the English and Polish exams, respectively, underscoring the inadequacy of standard MCQA scores as reliable indicators of genuine medical competence.

0 citationsRead paper
Recent publications

Latest Papers

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Jun 10, 2026

This study addresses the susceptibility of existing multiple-choice question answering (MCQA)–based evaluations of medical large language models to guessing and answer bias, which often inflate estimates of true clinical reasoning capabilities. To mitigate these limitations, the authors introduce a more rigorous benchmark based on Polish medical licensing examinations, incorporating over 15,000 new questions spanning two additional clinical domains and implementing four structural enhancements designed to reduce inherent MCQA biases. The benchmark facilitates cross-lingual evaluation and data contamination detection, and was used to systematically assess 21 prominent large language models. Results reveal that under this more challenging setting, the top-performing model, Qwen3.5-122B, exhibits performance drops of 28.4 and 31 percentage points on the English and Polish exams, respectively, underscoring the inadequacy of standard MCQA scores as reliable indicators of genuine medical competence.

0 citationsRead paper